Language Models are Few-Shot Learners
2020/05/28 by Tom B. Brown, T. B. Brown, Brown, Tom B. +61 · 16 voices · 3692 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2005.14165
openalex publication_date 2020/05/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.
Citations
Cited by
- Concept Concentration for Faithful Representation Intervention
- Gumbel Distillation for Parallel Text Generation
- Causal-AgentIR: Self-Evolving Causal Memory for Adaptive Image Restoration Agents
- Twins: Learn to Predict Unified Representations with Focal Loss
- What Matters When Building Universal Multilingual Named Entity Recognition Models?
- \kappa-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating
- Carpe Diem: Critical Learning Period-Aware Contract-Based Incentives for Federated Learning
- PoCEvolve: Generating Proof-of-Concept Exploits from Security Patches with Vulnerability-Aware Prompt Evolution
- Scaling Native Multimodal Pre-Training From Scratch
- Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes
- When Machines Lie Differently: Detecting AI vs Human Fake News
- Improving Large Vision-Language Models' Understanding for Flow Field Data
- Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging
- HarnessLLM: Rust Verification Harness Generation with Large Language Models
- From Grasping to Speaking: Generative AI-Based Environment-Grounded VR Communication Training for Autistic Individuals
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Pretraining Recurrent Networks without Recurrence
- Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints
- HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding
- Agentic Root Cause Analysis through Evidence-Grounded Reasoning
- Vibe Coding: An Experiment with Test-Driven Development
- A Unified Moral-Value Dataset for Instruction Tuning
- CRAFT: Exploring Wearable Creative AI on Smart Glasses for Fiction Writing in Real-World Contexts
- Stabilizing Native Low-Rank LLM Pretraining
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models
- Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance
- Efficient Clustering with Provable Guardrails for LLM Inference at Scale
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
- Grounding latent algorithm routing in transformer reasoning
- Towards Automated Formal Verification of zkEVMs Using LLM-Guided Constraint Synthesis
- Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens
- ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
- The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models
- Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
- Show Me Examples: Inferring Visual Concepts from Image Sets
- A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
- Post-Training in Time Series Foundation Models: A Unifying Framework
- LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
- In-span learning: adapting reduced-order models using their own predictions
- ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
- TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
- Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
- Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models
- HiCI: Hierarchical Construction-Integration for Long-Context Attention
- HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
- Stochastic Thermodynamics for Autoregressive Generative Models: A Non-Markovian Perspective
- VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems
- PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation
- Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
- Test-Time Registers as Global Priors for Tokenized Image Generation
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation
- Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection
- Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
- AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism
- Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
- Parameter-Efficient Continual Fine-Tuning: A Survey
- Auditing pretraining contamination in single-cell foundation model benchmarks
- DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
- LLM-Driven Cross-Paradigm Design for Quantum Optimal Control
- Tabular Foundation Models for Discrete Choice Estimation
- Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding
- Residual-Guided Multi-Resolution Refinement of Foundation Models: A Case Study in Drought Forecasting
- BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
- Planning with Transformers: Chain of Computation and Structured Context Windows
- MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
- What Transfers Under Source Shift? Definitions, Examples, and Fine-Tuning for Climate Disclosure Classification
- I wanted it to feel more personal: Customization of social AI as AI individualism in practice
- Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality
- Testing Retrieval-Augmented Generation Systems with Chunk Coverage
- Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding
- SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
- C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
- Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
- Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments
- Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
- Bigger Is Safer: Provable Robustness in In-Context Learning Scales with Capacity
- Counting Cycles with AI: Counting Cycles with AI: Computationally Efficient Equivalent Forms with Applications
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph
- Cognitive-YOLO: LLM-Driven Architecture Synthesis from First Principles of Data for Object Detection
- Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach
- ADEPT: Architecture-Driven Energy-Efficient CNN Fine-Tuning on PIM Accelerators
- Measuring and Evaluating the Performance of Generative AI Models for Scam Detection
- Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
- SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
- Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
- RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
- Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
- LogicIF: Towards Complex Logic Instruction Following
- Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
- Prompt-Guided Foundation Model Tuning for Pathology Image Classification
- Supervised Reward Inference
- InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
- NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs
- Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents
- Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
- Scaling Point-in-Time Language Models
- SportD: Can VLMs Physically Strategize?
- TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue
- Mapping the Narrow Corridor with Large Language Models
- An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
- I-Rex: An Interactive Debugger for SQL
- In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention
- Decoupled Alignment for Robust Plug-and-Play Adaptation
- ICLR: In-Context Imitation Learning with Visual Reasoning
- Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
- Knowledge-Centric Agents for Workflow Generation in ComfyUI
- Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative
- Revisiting data-driven dynamic security assessment with a tabular foundation model
- Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
- NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents
- Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry
- Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
- Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
- MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model
- EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
- AI-Conducted Interviews in Empirical Software Engineering: An Experience Report
- On-Policy Delta Distillation
- A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems
- Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
- Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
- CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA
- Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
- Mixtures of SubExperts for Large Language Continual Learning
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
- Robust Explanations for User Trust in Enterprise NLP Systems
- Data and trained models for "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication"
- Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
- HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization
- AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery
- Active Real-World Factor-Based Evaluation for Generalist Robot Policies
- ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
- Toward Robust In-Context Segmentation via Concept Guidance
- DiMaS: Distribution Matching for Steering Vision-Language-Action Models
- MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference
- How Can AI Augment Access to Justice? Public Defenders' Perspectives on Responsible AI Adoption
- Bayesian Wind Tunnels for Model Selection
- Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
- ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models
- Can In-Context Learning Support Intrinsic Curiosity?
- The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
- Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
- LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats
- From Errors to Rules: Iterative Prompt Optimization for Text Classification
- Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge Graphs
- Lifted Representation Hypothesis in Language Models
- Do Transformers Need Three Projections? Systematic Study of QKV Variants
- Automatically Attacking Software Reverse Engineering AI Agents
- Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
- The Hard Decision Layer: Evidence for Committed Inference in Transformers
- Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
- Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- AI-Mediated Communication Can Steer Collective Opinion
- LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets
- Leveraging Design-Aware Context in Large Language Models for Code Comment Generation
- Toward manifest relationality in transformers via symmetry reduction
- S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
- Response drift across frontier large language models
- A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction
- Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference
- GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
- 123D: Unifying Multi-Modal Autonomous Driving Data at Scale
- Accelerating Heterogeneous Agent Collaboration in Dynamic Edge Networks
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations
- T5-CSBoost: Adversarial Perturbation Resistant LLM Fingerprinting
- Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- UzWordnet and Generative AI for Learning Uzbek by Game Playing
- The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
- Improving LLM-Driven Test Generation by Learning from Mocking Information
- Stochasticity in Tokenisation Improves Robustness
- "AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
- Information bottleneck for learning the phase space of dynamics from high-dimensional experimental data
- Generalist AI control: Towards multi-purpose adaptive algorithms
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- StoryScope: Investigating idiosyncrasies in AI fiction
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context Learning
- Sparser, Faster, Lighter Transformer Language Models
- Cross-Modal Taxonomic Generalization in (Vision-) Language Models
- Position: Modular Memory is the Key to Continual Learning Agents
- Transformers for dynamical systems learn transfer operators in-context
- Mitigating Conversational Inertia in Multi-Turn Agents
- Path Integration and Object-Location Binding Emerge in an Action-Conditioned Predictive Sequence Network
- AI4SLT: Empirical Processes in Lean 4 for Formal Statistical Learning Theory
- Memory Caching: RNNs with Growing Memory
- Genomic perplexity and the evolution of context-dependent function
- Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High-Quality Books
- Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Learning Abstractions for Hierarchical Planning in Program-Synthesis Agents
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- It’s Not What You Say, It’s How You Say It: Evaluating LLM Responses to Expressions of Belief
- Linear representations in language models can change dramatically over a conversation
- Self-Distillation Enables Continual Learning
- How Human is AI? Examining the Impact of Emotional Prompts on Artificial and Human and Responsiveness
- mHC: Manifold-Constrained Hyper-Connections
- Shared sensitivity to data distribution during learning in humans and transformer networks
- NL2Logic: AST-Guided Translation of Natural Language into First-Order Logic with Large Language Models
- Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering
- RePo: Language Models with Context Re-Positioning
- Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
- A Network of Biologically Inspired Rectified Spectral Units (ReSUs) Learns Hierarchical Features Without Error Backpropagation
- Irresponsible AI: big tech's influence on AI research and associated impacts
- Epistemological Fault Lines Between Human and Artificial Intelligence
- A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
- Learning to Orchestrate Agents in Natural Language with the Conductor
- What does it mean to understand language?
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- The 4/δ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee
- Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- LLMs can hide text in other text of the same length
- SAM 3D: 3Dfy Anything in Images
- A Primer on Quantum Machine Learning
- Mind captioning: Evolving descriptive text of mental content from human brain activity
- Ten Simple Rules for AI-Assisted Coding in Science
- Butter-Bench: Evaluating LLM Controlled Robots for Practical Intelligence
- Jasmine: A Simple, Performant and Scalable JAX-based World Modeling Codebase
- AI use in American newspapers is widespread, uneven, and rarely disclosed
- Can We Hide Machines in the Crowd? Quantifying Equivalence in LLM-in-the-loop Annotation Tasks
- Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
- The Impossibility of Inverse Permutation Learning in Transformer Models
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- A Single Character can Make or Break Your LLM Evals
- Video models are zero-shot learners and reasoners
- Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
- Pre-training under infinite compute
- Towards a Physics Foundation Model
- Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
- Is In-Context Learning Learning?
- World Modeling with Probabilistic Structure Integration
- K2-Think: A Parameter-Efficient Reasoning System
- If generative AI is the answer, what is the question?
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
- Measuring Scalar Constructs in Social Science with LLMs
- Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
- Scaling language model size yields diminishing returns for single-message political persuasion
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
- Power Stabilization for AI Training Datacenters
- Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models
- A Survey on Diffusion Language Models
- Fast weight programming and linear transformers: from machine learning to neurobiology
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- Whither symbols in the era of advanced neural networks?
- Markov Chain Estimation with In-Context Learning
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- AlphaGo Moment for Model Architecture Discovery
- Measuring Negative Campaigning across Languages with Large Language Models: A Study of 18 Million Tweets in 19 Countries
- Learning without training: The implicit dynamics of in-context learning
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility
- Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
- SI-Agent: An Agentic Framework for Feedback-Driven Generation and Tuning of Human-Readable System Instructions for Large Language Models
- What Neuroscience Can Teach AI About Learning in Continuously Changing Environments
- Investigating age-related differences in semantic control mechanisms involved in creative cognition
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
- Fast and Simplex: 2-Simplicial Attention in Triton
- From Memories to Maps: Mechanisms of In-Context Reinforcement Learning in Transformers
- A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
- Simulating Society Requires Simulating Thought
- Text-to-LoRA: Instant Transformer Adaption
- Unsupervised pretraining in biological neural networks
- Mercury: Ultra-Fast Language Models Based on Diffusion
- Not All Tokens Are Meant to Be Forgotten
- Cortical language areas are coupled via a soft hierarchy of model-based linguistic features
- Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- General agents contain world models
- Self-supervised learning of molecular representations from millions of tandem mass spectra using DreaMS
- Distillation of atomistic foundation models across architectures and chemical domains
- Self-Adapting Language Models
- Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese
- Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning
- A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
- Relational reasoning and inductive bias in transformers and large language models
- Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
- Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
- Updating “The Future of Coding”: Qualitative Coding with Generative Large Language Models
- STORY2GAME: Generating (Almost) Everything in an Interactive Fiction Game
- 34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery
- El Agente: An autonomous agent for quantum chemistry
- True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics
- Of Mice and Machines: A Comparison of Learning Between Real World Mice and RL Agents
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
- Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- Type-Constrained Code Generation with Language Models
- Transfer between Modalities with MetaQueries
- Enough Coin Flips Can Make LLMs Act Bayesian
- Mixture of Experts Made Intrinsically Interpretable
- Trends in AI Supercomputers
- VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
- Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
- Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Do Chinese models speak Chinese languages?
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
- Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication
- VGGT: Visual Geometry Grounded Transformer
- Quantum advantage for learning shallow neural networks with natural data distributions
- Vision as LoRA
- SPECTRE: An FFT-Based Efficient Drop-In Replacement to Self-Attention for Long Contexts
- Large Language Diffusion Models
- Foundation neural-networks quantum states as a unified Ansatz for multiple hamiltonians
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data
- The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models
- TransMLA: Multi-Head Latent Attention Is All You Need
- Automated Capability Discovery via Foundation Model Self-Exploration
- Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
- On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks
- Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
- LM2: Large Memory Models
- Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
- Training Dynamics of In-Context Learning in Linear Attention
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
- Mirai: A Wearable Proactive AI "Inner-Voice" for Contextual Nudging
- The in-context inductive biases of vision-language models differ across modalities
- Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
- Evolution and The Knightian Blindspot of Machine Learning
- The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
- MedicoSAM: Robust Improvement of SAM for Medical Imaging
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- A Toolbox for Improving Evolutionary Prompt Search
- Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
- Connectomics Informed by Large Language Models
- Revisiting Theory of Contrastive Learning for Domain Generalization
- Exploring Depth Generalization in Large Language Models for Solving Recursive Logic Tasks
- Time Series Foundation Models for Process Model Forecasting
- OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- Do Large Language Models Truly Understand Cross-cultural Differences?
- It's LIT! Reliability-Optimized LLMs with Inspectable Tools
- An AI-Powered Autonomous Underwater System for Sea Exploration and Scientific Research
- From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change
- The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing
- Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
- Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
- Exploring the integration of large language models in industrial test maintenance processes
- A framework for assessing the capabilities of code generation of constraint domain-specific languages with large language models
- How Intrinsic Motivation Underlies Embodied Open-Ended Behavior
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- Reinforcement Learning via Self-Distillation
- PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image Generation
- Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects?
- A Stepwise-Enhanced Reasoning Framework for Large Language Models Based on External Subgraph Generation
- Towards end-to-end automation of AI research
- MedGemma vs GPT-4: Open-Source and Proprietary Zero-shot Medical Disease Classification from Images
- Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control
- Diffusion-based Decentralized Federated Multi-Task Representation Learning
- Taxing Artificial Intelligence
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Beyond Cosine Similarity
- Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Multiple Token Divergence: Measuring and Steering In-Context Computation Density
- Efficiency vs. understanding: a critical examination of ChatGPT’s performance in context-sensitive annotation tasks
- Eliciting Behaviors in Multi-Turn Conversations
- Debugging Tabular Log as Dynamic Graphs
- DECEPTICON: How Dark Patterns Manipulate Web Agents
- Harnessing Large Language Models for Biomedical Named Entity Recognition
- Beyond Centralization: Provable Communication Efficient Decentralized Multi-Task Learning
- When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
- Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis
- Learning When Not to Attend Globally
- Role-Based Fault Tolerance System for LLM RL Post-Training
- Towards Efficient Post-Training via Fourier-Driven Adapter Architectures
- DeMoGen: Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
- GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
- HELIOS: An LLM-Driven Autonomous Indirect Trajectory Optimization Agent
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
- Understanding Tone-Dependent Inference Cost in Large Language Models
- Greedy dynamical meta-learning
- Distributed Sparse Interventions in Language Models
- The Weaponization of Computer Vision: Tracing Military-Surveillance Ties through Conference Sponsorship
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
- Verification-Notebook Learning for Source-Aware Multimodal Misinformation Detection
- seqLens: Optimizing Language Models for Genomic Predictions
- BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi
- Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Enhancing Code Understanding for Impact Analysis by Combining Transformers and Program Dependence Graphs
- Training Language Models to Cooperate with Inference-Time Controllers
- Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
- Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails
- Generative AI for Requirements Engineering: A Systematic Literature Review
- Generative Artificial Intelligence for Software Engineering -- A Research Agenda
- The Impact of Prompt Programming on Function-Level Code Generation
- A Scoping Review of Machine Learning Applications in Power System Protection and Disturbance Management
- SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation
- Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation
- DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems
- The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation
- Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
- PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
- Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch
- LLMs for Agentic Home Energy Management
- VaLiDRec: Variable-Length LLM-Aligned Semantic IDs for Generative Recommendation
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
- LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
- In-Context Learning as Implicit Policy Gradient
- A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
- When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness
- Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM-Based Product Attribute Extraction
- Language-Routed RAG and Direct Option Scoring for Multilingual Financial QA: DS@GT at FinMMEval
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests
- Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- HiLLTS: Zero-Shot Hierarchical LLM-Guided Traffic Signal Control for Sustainable Transportation
- Imprompt: A Language Framework for Prompt Programming
- Influence of Prompt Engineering on Small Language Models for Guarded Query Routing
- Toward Human-Centered Multi-Agent Systems: Integrating Cognition, Culture, Values, and Cooperation in AI Agents
- PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF
- CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
- TRE: Training-Free Hallucination Detection for Diffusion Language Models
- ARdena: Scenario-driven control of real-time LLM agents
- Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
- cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs
- Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- The Cartesian Cut in Agentic AI
- Adapting a Text-to-Audio Model for Room Impulse Response Generation
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- Cognitive Dark Matter: Measuring What AI Misses
- Explainability Methods for Hardware Trojan Detection: A Systematic Comparison
- LVLM-Aided Alignment of Task-Specific Vision Models
- Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
- CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
- TimeBill: Time-Budgeted Inference for Large Language Models
- Knowledge Reasoning of Large Language Models Integrating Graph-Structured Information for Pest and Disease Control in Tobacco
- Method Decoration (DeMe): A Framework for LLM-Driven Adaptive Method Generation in Dynamic IoT Environments
- Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
- Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
- Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant
- Hierarchy-Aware Fine-Tuning of Vision-Language Models
- Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
- The AI Committee: A Multi-Agent Framework for Automated Validation and Remediation of Web-Sourced Data
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
- Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning
- Dynamic Attention (DynAttn): Interpretable High-Dimensional Spatio-Temporal Forecasting (with Application to Conflict Fatalities)
- What Makes a GitHub Issue Ready for Copilot?
- Parallel Token Prediction for Language Models
- Beyond Context: Large Language Models Failure to Grasp Users Intent
- Uncovering Hierarchical Structure in LLM Embeddings with δ-Hyperbolicity, Ultrametricity, and Neighbor Joining
- Understanding Scaling Laws in Deep Neural Networks via Feature Learning Dynamics
- When LLMs fall short in Deductive Coding: Model Comparison and Human AI Collaboration Workflow Design
- Artificial or Just Artful? Do LLMs Bend the Rules in Programming?
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- FinAgent: An Agentic AI Framework Integrating Personal Finance and Nutrition Planning
- Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
- Pioneering Multimodal Emotion Recognition in the Era of Large Models: From Closed Sets to Open Vocabularies
- Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
- Assessing the Software Security Comprehension of Large Language Models
- Semantic Deception: When Reasoning Models Can't Compute an Addition
- FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- Neural Scaling Laws for Learning-based Identification of Nonlinear Systems
- Enhancing Zero-Shot Time Series Forecasting in Off-the-Shelf LLMs via Noise Injection
- SynCraft: Guiding Large Language Models to Predict Edit Sequences for Molecular Synthesizability Optimization
- Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations
- LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs
- Adaptive Financial Sentiment Analysis for NIFTY 50 via Instruction-Tuned LLMs , RAG and Reinforcement Learning Approaches
- LoFT-LLM: Low-Frequency Time-Series Forecasting with Large Language Models
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
- Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
- Fine-Tuned In-Context Learners for Efficient Adaptation
- Vehicle-centric Perception via Multimodal Structured Pre-training
- Attention Is Not What You Need
- Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
- Exploring the features used for summary evaluation by Human and GPT
- Increasing the Thinking Budget is Not All You Need
- Event Extraction in Large Language Model
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- The Mental World of Large Language Models in Recommendation: A Benchmark on Association, Personalization, and Knowledgeability
- CoDrone: Autonomous Drone Navigation Assisted by Edge and Cloud Foundation Models
- Can abstract concepts from LLM improve SLM performance?
- Efficient Personalization of Generative Models via Optimal Experimental Design
- DREAM: Dynamic Red-teaming across Environments for AI Models
- R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression
- FASTRIC: Prompt Specification Language for Verifiable LLM Interactions
- Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- Beyond the Prompt: An Empirical Study of Cursor Rules
- MDToC: Metacognitive Dynamic Tree of Concepts for Boosting Mathematical Problem-Solving of Large Language Models
- HARBOR: Holistic Adaptive Risk assessment model for BehaviORal healthcare
- A Study of Finetuning Video Transformers for Multi-view Geometry Tasks
- ASTIF: Adaptive Semantic-Temporal Integration for Cryptocurrency Price Forecasting
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
- A Comparative Study of Light-weight Language Models for PII Masking and their Deployment for Real Conversational Texts
- Reflective Confidence: Correcting Reasoning Flaws via Online Self-Correction
- MoE Pathfinder: Trajectory-driven Expert Pruning
- AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
- TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
- A Data-Centric Approach to Generalizable Speech Deepfake Detection
- AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3
- Fairness Is Not Just Ethical: Performance Trade-Off via Data Correlation Tuning to Mitigate Bias in ML Software
- Shuttling Compiler for Trapped-Ion Quantum Computers Based on Large Language Models
- Adversarial Robustness of Vision in Open Foundation Models
- GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
- Mitty: Diffusion-based Human-to-Robot Video Generation
- Task Schema and Binding: A Double Dissociation Study of In-Context Learning
- Stakeholder Suite: A Unified AI Framework for Mapping Actors, Topics and Arguments in Public Debates
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- TCDE: Topic-Centric Dual Expansion of Queries and Documents with Large Language Models for Information Retrieval
- LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering
- Next-Embedding Prediction Makes Strong Vision Learners
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- In-Context Algebra
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- Meta-RL Induces Exploration in Language Agents
- LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
- NRGPT: An Energy-based Alternative for GPT
- Abacus: Self-Supervised Event Counting-Aligned Distributional Pretraining for Sequential User Modeling
- A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection
- Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
- Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
- Emergent Bias and Fairness in Multi-Agent Decision Systems
- Introducing ORKG ASK: an AI-driven Scholarly Literature Search and Exploration System Taking a Neuro-Symbolic Approach
- Hypernetworks That Evolve Themselves
- Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
- Empirical Likelihood Meets Prediction-Powered Inference
- Evaluating OpenAI GPT Models for Translation of Endangered Uralic Languages: A Comparison of Reasoning and Non-Reasoning Architectures
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
- In-Context Multi-Operator Learning with DeepOSets
- ContextLeak: Auditing Leakage in Private In-Context Learning Methods
- Topic Discovery and Classification for Responsible Generative AI Adaptation in Higher Education
- Higher-Order LaSDI: Reduced Order Modeling with Multiple Time Derivatives
- BRAID: Bounded Reasoning for Autonomous Inference and Decisions
- AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
- In-Context Semi-Supervised Learning
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- Topological Metric for Unsupervised Embedding Quality Evaluation
- IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
- Bolmo: Byteifying the Next Generation of Language Models
- Reducing Pilots in Channel Estimation with Predictive Foundation Models
- An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
- Case Prompting to Mitigate Large Language Model Bias for ICU Mortality Prediction
- ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata
- Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
- Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- The Semantic Architect: How FEAML Bridges Structured Data and LLMs for Multi-Label Tasks
- An Exploratory Study of Bayesian Prompt Optimization for Test-Driven Code Generation with Large Language Models
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- SeBERTis: A Framework for Producing Classifiers of Security-Related Issue Reports
- DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
- PPSEBM: An Energy-Based Model with Progressive Parameter Selection for Continual Learning
- Quantifying Return on Security Controls in LLM Systems
- Few-Shot Inference of Human Perceptions of Robot Performance in Social Navigation Scenarios
- Spatia: Video Generation with Updatable Spatial Memory
- Evaluating Code Reasoning Abilities of Large Language Models Under Real-World Settings
- Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- Spherical Leech Quantization for Visual Tokenization and Generation
- Let the Barbarians In: How AI Can Accelerate Systems Performance Research
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
- Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
- Dual-objective Language Models: Training Efficiency Without Overfitting
- RecGPT-V2 Technical Report
- C-ing Clearly: Enhanced Binary Code Explanations using C code
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- Attention-Based Foundation Model for Quantum States
- Inflation Attitudes of Large Language Models
- TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
- LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
- CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- Machine learning discovers new champion codes
- Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
- Towards Interactive Intelligence for Digital Humans
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- Scaling Laws for Code: Every Programming Language Matters
- Improving Recursive Transformers with Mixture of LoRAs
- From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
- Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection
- CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs
- On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
- Sliding Window Recurrences for Sequence Models
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- Resting Neurons, Active Insights: Improving Input Sparsification for Large Language Models
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
- MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
- ORIBA: Exploring LLM-Driven Role-Play Chatbot as a Creativity Support Tool for Original Character Artists
- Human-Inspired Learning for Large Language Models via Obvious Record and Maximum-Entropy Method Discovery
- Content-Aware Ad Banner Layout Generation with Two-Stage Chain-of-Thought in Vision Language Models
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
- Low-Rank Compression of Language Models via Differentiable Rank Selection
- DL3M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models
- Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection
- WATOS: Efficient LLM Training Strategies and Architecture Co-exploration for Wafer-scale Chip
- Semantic Distance Measurement based on Multi-Kernel Gaussian Processes
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
- VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
- Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
- Bridging Streaming Continual Learning via In-Context Large Tabular Models
- The Effect of Document Summarization on LLM-Based Relevance Judgments
- AI Benchmark Democratization and Carpentry
- In-Context Learning for Seismic Data Processing
- Fully Inductive Node Representation Learning via Graph View Transformation
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- LOOPRAG: Enhancing Loop Transformation Optimization with Retrieval-Augmented Large Language Models
- DynaPURLS: Dynamic Refinement of Part-aware Representations for Skeleton-based Zero-Shot Action Recognition
- Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
- AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
- A Simple Generalisation of the Implicit Dynamics of In-Context Learning
- ReactorFold: Generative discovery of nuclear reactor cores via emergent physical reasoning
- amc: The Automated Mission Classifier for Telescope Bibliographies
- BAID: A Benchmark for Bias Assessment of AI Detectors
- Natural Language Interaction for Editing Visual Knowledge Graphs
- AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
- Fairness-Regularized Online Optimization with Switching Costs
- Your plan may succeed, but what about failure? Investigating how people use ChatGPT for long-term life task planning
- KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- Multi-Granular Node Pruning for Causal Circuit Discovery
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- LabelFusion: Learning to Fuse LLMs and Transformer Classifiers for Robust Text Classification
- Natural Language Interface for Firewall Configuration
- CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator
- LLM-Auction: Generative Auction towards LLM-Native Advertising
- Zero-shot 3D Map Generation with LLM Agents: A Dual-Agent Architecture for Procedural Content Generation
- Translating Informal Proofs into Formal Proofs Using a Chain of States
- Linear socio-demographic representations emerge in Large Language Models from indirect cues
- Reverse Thinking Enhances Missing Information Detection in Large Language Models
- MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation
- Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis
- Watermarks for Language Models via Probabilistic Automata
- Grounding Everything in Tokens for Multimodal Large Language Models
- Enhancing Next-Generation Language Models with Knowledge Graphs: Extending Claude, Mistral IA, and GPT-4 via KG-BERT
- Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
- Token Sample Complexity of Attention
- PARAN: Persona-Augmented Review ANswering system on Food Delivery Review Dataset
- Generate-Then-Validate: A Novel Question Generation Approach Using Small Language Models
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- SEMDICE: Off-policy State Entropy Maximization via Stationary Distribution Correction Estimation
- DynaMate: An Autonomous Agent for Protein-Ligand Molecular Dynamics Simulations
- Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
- SCOPE: Language Models as One-Time Teacher for Hierarchical Planning in Text Environments
- FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning
- Defining Cost Function of Steganography with Large Language Models
- Neurosymbolic Information Extraction from Transactional Documents
- LogICL: Distilling LLM Reasoning to Bridge the Semantic Gap in Cross-Domain Log Anomaly Detection
- Chasing Shadows: Pitfalls in LLM Security Research
- Supporting Dynamic Agentic Workloads: How Data and Agents Interact
- ODMA: On-Demand Memory Allocation Framework for LLM Serving on LPDDR-Class Accelerators
- Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
- Impact of Positional Encoding: Clean and Adversarial Rademacher Complexity for Transformers under In-Context Regression
- Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval
- Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts
- Can TabPFN Compete with GNNs for Node Classification via Graph Tabularization?
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
- Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot Learning
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
- Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
- MobileFineTuner: A Unified End-to-End Framework for Fine-Tuning LLMs on Mobile Phones
- Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
- Universal Adversarial Suffixes Using Calibrated Gumbel-Softmax Relaxation
- Learning and Editing Universal Graph Prompt Tuning via Reinforcement Learning
- Bridging Scale Discrepancies in Robotic Control via Language-Based Action Representations
- SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- Provable Long-Range Benefits of Next-Token Prediction
- In-Context and Few-Shots Learning for Forecasting Time Series Data based on Large Language Models
- Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
- AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- Training Language Models to Use Prolog as a Tool
- ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation
- Materium: An Autoregressive Approach for Material Generation
- LLM Use for Mental Health: Crowdsourcing Users' Sentiment-based Perspectives and Values from Social Discussions
- Exploiting the Randomness of Large Language Models (LLM) in Text Classification Tasks: Locating Privileged Documents in Legal Matters
- Procrustean Bed for AI-Driven Retrosynthesis: A Unified Framework for Reproducible Evaluation
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
- FOAM: Blocked State Folding for Memory-Efficient LLM Training
- Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph
- Block Sparse Flash Attention
- Bita: A Conversational Assistant for Fairness Testing
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- Large Language Model-Based Generation of Discharge Summaries
- Optimal and Diffusion Transports in Machine Learning
- A Patient-Doctor-NLP-System to contest inequality for less privileged
- Stochasticity in Agentic Evaluations: Quantifying Inconsistency with Intraclass Correlation
- Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
- Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
- Hybrid Quantum-Classical Ensemble Learning for S&P 500 Directional Prediction
- Small Language Models Can Use Nuanced Reasoning For Health Science Research Classification: A Microbial-Oncogenesis Case Study
- Efficient Text Classification with Conformal In-Context Learning
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- GENIUS: An Agentic AI Framework for Autonomous Design and Execution of Simulation Protocols
- AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- LOCUS: A System and Method for Low-Cost Customization for Universal Specialization
- Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety
- Metaphor-based Jailbreak Attacks on Text-to-Image Models
- Empathy by Design: Aligning Large Language Models for Healthcare Dialogue
- Compass: Co-Exploration of Mapping and Hardware for Heterogeneous Multi-Chiplet Accelerators Targeting LLM Inference Service Workloads
- From Text to Returns: Using Large Language Models for Mutual Fund Portfolio Optimization and Risk-Adjusted Allocation
- Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
- LA-RL: Language Action-guided Reinforcement Learning with Safety Guarantees for Autonomous Highway Driving
- The Road of Adaptive AI for Precision in Cybersecurity
- Credal and Interval Deep Evidential Classifications
- AI & Human Co-Improvement for Safer Co-Superintelligence
- Structured Document Translation via Format Reinforcement Learning
- AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
- Bridging Traditional Machine Learning and Large Language Models: A Two-Part Course Design for Modern AI Education
- STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
- Are LLMs Truly Multilingual? Exploring Zero-Shot Multilingual Capability of LLMs for Information Retrieval: An Italian Healthcare Use Case
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- Qwen3.5-Omni Technical Report
- Eval Factsheets: A Structured Framework for Documenting AI Evaluations
- Norm-Governed Multi-Agent Decision-Making in Simulator-Coupled Environments:The Reinsurance Constrained Multi-Agent Simulation Process (R-CMASP)
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- Towards an AI Fluid Scientist: LLM-Powered Scientific Discovery in Experimental Fluid Mechanics
- Personalizing Agent Privacy Decisions via Logical Entailment
- LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models
- Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
- Solving LLM Repetition Problem in Production: A Comprehensive Study of Multiple Solutions
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
- The Initialization Determines Whether In-Context Learning Is Gradient Descent
- UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
- Empirical Prompt Engineering for Construct Identification with Large Language Models
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- MANTRA: a Framework for Multi-stage Adaptive Noise TReAtment During Training
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- SRPG: Semantically Reconstructed Privacy Guard for Zero-Trust Privacy in Educational Multi-Agent Systems
- MemVerse: Multimodal Memory for Lifelong Learning Agents
- Data-Free Pruning of Self-Attention Layers in LLMs
- AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- YOLOA: Real-Time Affordance Detection via LLM Adapter
- Nexus: Higher-Order Attention Mechanisms in Transformers
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- Epistemic Substitution: How Grokipedia's AI-Generated Encyclopedia Restructures Authority
- The Homological Brain: Parity Principle and Amortized Inference
- Fairness-Aware Fine-Tuning of Vision-Language Models for Medical Glaucoma Diagnosis
- PretrainZero: Reinforcement Active Pretraining
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
- TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
- In-Context Sync-LoRA for Portrait Video Editing
- Lumos: Let there be Language Model System Certification
- GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
- Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules
- Benchmarking machine learning models for multi-class state recognition in double quantum dot data
- Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
- CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
- Q-BERT4Rec: Quantized Semantic-ID Representation Learning for Multimodal Recommendation
- UniCom: Towards a Unified and Cohesiveness-aware Framework for Community Search and Detection
- Offloading Artificial Intelligence Workloads across the Computing Continuum by means of Active Storage Systems
- promptolution: A Unified, Modular Framework for Prompt Optimization
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- Progressive Image Restoration via Text-Conditioned Video Generation
- Parameter-Efficient Subspace Optimization for LLM Fine-Tuning
- LLM-Driven Corrective Robot Operation Code Generation with Static Text-Based Simulation
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- Improving Phishing Resilience with AI-Generated Training: Evidence on Prompting, Personalization, and Duration
- Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
- Cross-Lingual Interleaving for Speech Language Models
- BrepGPT: Autoregressive B-rep Generation with Voronoi Half-Patch
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- Evaluating SAM2 for Video Semantic Segmentation
- Generating REST API Tests With Descriptive Names
- In-Context Learning for Deep Joint Source-Channel Coding Over MIMO Channels
- Two-Dimensional Quantization for Geometry-Aware Audio Coding
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- Securing Large Language Models (LLMs) from Prompt Injection Attacks
- polyRETRO: a Language Model Approach to predict Polymerization Class and Monomer(s) for a Target Polymer
- TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
- From Regression to Classification: Exploring the Benefits of Categorical Representations of Energy in MLIPs
- A Knowledge-Based Language Model: Deducing Grammatical Knowledge in a Multi-Agent Language Acquisition Simulation
- Zero-Training Temporal Drift Detection for Transformer Sentiment Models: A Comprehensive Analysis on Authentic Social Media Streams
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation
- REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
- UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
- SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- 3D-Consistent Multi-View Editing by Correspondence Guidance
- Simplex-Optimized Hybrid Ensemble for Large Language Model Text Detection Under Generative Distribution Drif
- Aligning Probabilistic Beliefs under Informative Missingness: LLM Steerability in Clinical Reasoning
- Cross-Geometry Transfer Learning in Fast Electromagnetic Shower Simulation
- Framework-Aware Code Generation with API Knowledge Graph-Constructed Data: A Study on HarmonyOS
- Progressive Code Integration for Abstractive Bug Report Summarization
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education
- Teleportation-Based Defenses for Privacy in Approximate Machine Unlearning
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Quantized-Tinyllava: a new multimodal foundation model enables efficient split learning
- Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
- SimScale: Learning to Drive via Real-World Simulation at Scale
- The Geometry of Certainty: Recursive Topological Condensation and the Limits of Inference
- Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- Listwise Preference Optimization with Element-wise Confusions for Aspect Sentiment Quad Prediction
- An Empirical Study on the Security Vulnerabilities of GPTs
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
- Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
- Masked Diffusion for Generative Recommendation
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Experts are all you need: A Composable Framework for Large Language Model Inference
- AgentShield: Make MAS more secure and efficient
- A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
- JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
- HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
- Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
- Artwork Interpretation with Vision Language Models: A Case Study on Emotions and Emotion Symbols
- Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
- Group-Aware Partial Model Merging for Children's Automatic Speech Recognition
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
- Automated Generation of MDPs Using Logic Programming and LLMs for Robotic Applications
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- TAGFN: A Text-Attributed Graph Dataset for Fake News Detection in the Age of LLMs
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- ABounD: Adversarial Boundary-Driven Few-Shot Learning for Multi-Class Anomaly Detection
- FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
- Unexplored flaws in multiple-choice VQA evaluations
- Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation
- UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
- Tacit Bidder-Side Collusion: Artificial Intelligence in Dynamic Auctions
- Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Bootstrapping LLMs via Preference-Based Policy Optimization
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- Merge and Bound: Direct Manipulations on Weights for Class Incremental Learning
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- Learning Multi-Order Block Structure in Higher-Order Networks
- Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
- AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
- Maglev-Pentabot: Magnetic Levitation System for Non-Contact Manipulation using Deep Reinforcement Learning
- Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
- Towards Audio Token Compression in Large Audio Language Models
- CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
- Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
- Reinforcement Learning for Latent-Space Thinking in LLMs
- On the Origin of Algorithmic Progress in AI
- CarBench: A Comprehensive Benchmark for Neural Surrogates on High-Fidelity 3D Car Aerodynamics
- Structured Prompting Enables More Robust Evaluation of Language Models
- Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
- Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
- Image2Gcode: Image-to-G-code Generation for Additive Manufacturing Using Diffusion-Transformer Model
- ROOT: Robust Orthogonalized Optimizer for Neural Network Training
- Translating Large-Scale C Repositories to Idiomatic Rust
- Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection
- Tiny-TSM: Efficiently Training a Lightweight SOTA Time Series Foundation Model
- From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding
- Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Generation, Evaluation, and Explanation of Novelists' Styles with Single-Token Prompts
- Block Cascading: Training Free Acceleration of Block-Causal Video Models
- LLMs for Automated Unit Test Generation and Assessment in Java: The AgoneTest Framework
- Geometry of Decision Making in Language Models
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- LLM-Driven Transient Stability Assessment: From Automated Simulation to Neural Architecture Design
- Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
- CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
- Explainable Visual Anomaly Detection via Concept Bottleneck Models
- More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
- Foundry: Distilling 3D Foundation Models for the Edge
- ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- Annotation-Free Class-Incremental Learning
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
- Schema Matching on Graph: Iterative Graph Exploration for Efficient and Explainable Data Integration
- The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models
- TREASURE: The Visa Payment Foundation Model for High-Volume Transaction Understanding
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
- HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization
- Pretraining Transformer-Based Models on Diffusion-Generated Synthetic Graphs for Alzheimer's Disease Prediction
- HunyuanOCR Technical Report
- Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Learning Plug-and-play Memory for Guiding Video Diffusion Models
- Can LLMs Threaten Human Survival? Benchmarking Potential Existential Threats from LLMs via Prefix Completion
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- LLMs-Powered Real-Time Fault Injection: An Approach Toward Intelligent Fault Test Cases Generation
- A Multi-Agent LLM Framework for Multi-Domain Low-Resource In-Context NER via Knowledge Retrieval, Disambiguation and Reflective Analysis
- A Longitudinal Measurement of Privacy Policy Evolution for Large Language Models
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Emotion-Aware Conversational Recommender Systems: a Case Study
- SmartPoC: Generating Executable and Validated PoCs for Smart Contract Bug Reports
- Optimizing LLM Code Suggestions: Feedback-Driven Timing with Lightweight State Bounds
- Concept than Document: Context Compression via AMR-based Conceptual Entropy
- Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
- Scale What Counts, Mask What Matters: Evaluating Foundation Models for Zero-Shot Cross-Domain Wi-Fi Sensing
- Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- Empathetic Cascading Networks: A Multi-Stage Prompting Technique for Reducing Social Biases in Large Language Models
- LLMAID: Identifying AI Capabilities in Android Apps with LLMs
- IRSDA: An Agent-Orchestrated Framework for Enterprise Intrusion Response
- Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
- Prompt Optimization as a State-Space Search Problem
- Semantics as a Shield: Label Disguise Defense (LDD) against Prompt Injection in LLM Sentiment Classification
- Leveraging Language Models for Interpretable Analysis of Narratives in a Large Corpus
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- A Systematic Study of Compression Ordering for Large Language Models
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- A2Flow: Automating Agentic Workflow Generation via Self-Adaptive Abstraction Operators
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
- DiscoVerse: Multi-Agent Pharmaceutical Co-Scientist for Traceable Drug Discovery and Reverse Translation
- Efficient Inference Using Large Language Models with Limited Human Data: Fine-Tuning then Rectification
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- Towards Automating Data Access Permissions in AI Agents
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
- Controllability Analysis of State Space-based Language Model
- Synthesizing Precise Protocol Specs from Natural Language for Effective Test Generation
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
- MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
- AI- and Ontology-Based Enhancements to FMEA for Advanced Systems Engineering: Current Developments and Future Directions
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- Large Language Models for Sentiment Analysis to Detect Social Challenges: A Use Case with South African Languages
- Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- Efficient Robot Design with Multi-Objective Black-Box Optimization and Large Language Models
- A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- Spanning Tree Autoregressive Visual Generation
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- MedImageInsight for Thoracic Cavity Health Classification from Chest X-rays
- Affective Multimodal Agents with Proactive Knowledge Grounding for Emotionally Aligned Marketing Dialogue
- Bridging Symbolic Control and Neural Reasoning in LLM Agents: The Structured Cognitive Loop
- Fine-grained MoE Load Balancing with Linear Programming
- Chatbots to strengthen democracy: An interdisciplinary seminar to train identifying argumentation techniques of science denial
- Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
- Integrating Symbolic Natural Language Understanding and Language Models for Word Sense Disambiguation
- Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation
- AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
- ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
- SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
- Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
- PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
- Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
- AskDB: An LLM Agent for Natural Language Interaction with Relational Databases
- BlockCert: Certified Blockwise Extraction of Transformer Mechanisms
- ILoRA: Federated Learning with Low-Rank Adaptation for Heterogeneous Client Aggregation
- TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues
- KRAL: Knowledge and Reasoning Augmented Learning for LLM-assisted Clinical Antimicrobial Therapy
- JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
- iLTM: Integrated Large Tabular Model
- SpectralTrain: A Universal Framework for Hyperspectral Image Classification
- CARE-RAG - Clinical Assessment and Reasoning in RAG
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
- RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue
- Fault2Flow: An AlphaEvolve-Optimized Human-in-the-Loop Multi-Agent System for Fault-to-Workflow Automation
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Towards Understanding Layer Contributions in Tabular In-Context Learning Models
- Automatic Pruning Discovery for Large Language Models
- Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
- DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
- Reflexive Evidence-Based Multimodal Learning for Clean Energy Transitions: Causal Insights on Cooking Fuel Access, Urbanization, and Carbon Emissions
- An Optimized Machine Learning Classifier for Detecting Fake Reviews Using Extracted Features
- As If We've Met Before: LLMs Exhibit Certainty in Recognizing Seen Files
- SafeRBench: A Comprehensive Benchmark for Safety Assessment in Large Reasoning Models
- Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- DEPO: Dual-Efficiency Preference Optimization for LLM Agents
- GPS: General Per-Sample Prompter
- Strategic Innovation Management in the Age of Large Language Models Market Intelligence, Adaptive R&D, and Ethical Governance
- M-CALLM: Multi-level Context Aware LLM Framework for Group Interaction Prediction
- Gradient-descent methods for quantum detector tomography
- LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
- Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
- Mitigating Label Length Bias in Large Language Models
- ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
- Report on the Scoping Workshop on AI in Science Education Research 2025
- AI-driven Generation of MALDI-TOF MS for Microbial Characterization
- CascadedViT: Cascaded Chunk-FeedForward and Cascaded Group Attention Vision Transformer
- Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
- TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- Node-Level Uncertainty Estimation in LLM-Generated SQL
- Personality Pairing Improves Human-AI Collaboration
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
- Applying Large Language Models to Characterize Public Narratives
- SnapAudit: Active Auditing of Differentially Private In-Context Learning via Snapshot-Based Simulation
- Multi-Agent Multimodal Large Language Model Framework for Automated Interpretation of Fuel Efficiency Analytics in Public Transportation
- Donors and Recipients: On Asymmetric Transfer Across Tasks and Languages with Parameter-Efficient Fine-Tuning
- Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
- Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework
- Zero-Shot Grammar Competency Estimation Using Large Language Model Generated Pseudo Labels
- Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
- Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
- Knowing Ourselves Through Others: Reflecting with AI in Digital Human Debates
- SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
- Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- Online Continual Learning on Intel Loihi 2 via a Co-designed Spiking Neural Network
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- Multimodal Large Language Models as Image Classifiers
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- Evidence of Phase Transitions in Small Transformer-Based Language Models
- Prompt-Driven Domain Adaptation for End-to-End Autonomous Driving via In-Context RL
- AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
- Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
- Designed to Spread: Generative Approaches to Enhance Information Diffusion
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
- 47B Mixture-of-Experts Beats 671B Dense Models on Chinese Medical Examinations
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Prompt Engineering Techniques for Context-dependent Text-to-SQL in Arabic
- Genomic Next-Token Predictors are In-Context Learners
- Mitigating Length Bias in RLHF through a Causal Lens
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Adaptive Focus Memory for Language Models
- GenSIaC: Toward Security-Aware Infrastructure-as-Code Generation with Large Language Models
- BitSnap: Checkpoint Sparsification and Quantization in LLM Training
- AstroMLab 5: Structured Summaries and Concept Extraction for 400,000 Astrophysics Papers
- Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
- Debate over Mixed-knowledge: A Robust Multi-Agent Reasoning Framework for Incomplete Knowledge Graph Question Answering
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?
- MixAR: Mixture Autoregressive Image Generation
- OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- Leveraging Large Language Models for Career Mobility Analysis: A Study of Gender, Race, and Job Change Using U.S. Online Resume Profiles
- Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning
- A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
- Defending Unauthorized Model Merging via Dual-Stage Weight Protection
- Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
- On the Notion that Language Models Reason
- CURENet: Combining Unified Representations for Efficient Chronic Disease Prediction
- Free3D: 3D Human Motion Emerges from Single-View 2D Supervision
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text
- LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Hindsight Distillation Reasoning with Knowledge Encouragement Preference for Knowledge-based Visual Question Answering
- AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
- Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- Generative Caching for Structurally Similar Prompts and Responses
- Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
- On the Measure of a Model: From Intelligence to Generality
- Socrates-Mol: Self-Oriented Cognitive Reasoning through Autonomous Trial-and-Error with Empirical-Bayesian Screening for Molecules
- Belief Net: A Filter-Based Framework for Learning Hidden Markov Models from Observations
- POTSA: A Cross-Lingual Speech Alignment Framework for Low Resource Speech-to-Text Translation
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- Beyond Elicitation: Provision-based Prompt Optimization for Knowledge-Intensive Tasks
- Analogical Structure, Minimal Contextual Cues and Contrastive Distractors: Input Design for Sample-Efficient Linguistic Rule Induction
- RoboBenchMart: Benchmarking Robots in Retail Environment
- ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
- Persona-Aware Alignment Framework for Personalized Dialogue Generation
- Physics-informed Machine Learning for Static Friction Modeling in Robotic Manipulators Based on Kolmogorov-Arnold Networks
- Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
- How Can We Effectively Use LLMs for Phishing Detection?: Evaluating the Effectiveness of Large Language Model-based Phishing Detection Models
- NumPert: Numerical Perturbations to Probe Language Models for Veracity Prediction
- Boosting In-Silicon Directed Evolution with Fine-Tuned Protein Language Model and Tree Search
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
- From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models
- Mastering Olympiad-Level Physics with Artificial Intelligence
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- Context-Aware Multimodal Representation Learning for Spatio-Temporally Explicit Environmental Modelling
- Transformer Semantic Genetic Programming for d-dimensional Symbolic Regression Problems
- BarrierBench : Evaluating Large Language Models for Safety Verification in Dynamical Systems
- Multi-step Predictive Coding Leads To Simplicity Bias
- Review of Passenger Flow Modelling Approaches Based on a Bibliometric Analysis
- Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
- Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
- The Path Not Taken: RLVR Provably Learns Off the Principals
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
- | \circlearrowright \boxedBUS |: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
- Leveraging unlabelled data for generalizable neural population decoding
- MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
- Concentration bounds on response-based vector embeddings of black-box generative models
- ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation
- Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
- Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
- VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- EHRStruct: A Comprehensive Benchmark Framework for Evaluating Large Language Models on Structured Electronic Health Record Tasks
- DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
- Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
- Relation as a Prior: A Novel Paradigm for LLM-based Document-level Relation Extraction
- PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
- Generalizable Insights for Graph Transformers in Theory and Practice
- Digital Nature Revisited: A Ten-Year Synthesis of Art, Technology, and the Evolution of "Nature": Reimagining Post-Truth Ecologies Through Art, Algorithm, and Animism
- NOTAM-Evolve: A Knowledge-Guided Self-Evolving Optimization Framework with LLMs for NOTAM Interpretation
- Accelerating Training Speed of Tiny Recursive Models with Curriculum Guided Adaptive Recursion
- Multi-objective Hyperparameter Optimization in the Age of Deep Learning
- Neurophysiological Characteristics of Adaptive Reasoning for Creative Problem-Solving Strategy
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
- CellARC: Measuring Intelligence with Cellular Automata
- Data Descriptions from Large Language Models with Influence Estimation
- Parallel Sampling via Autospeculation
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- Majority Rules: LLM Ensemble is a Winning Approach for Content Categorization
- CC30k: A Citation Contexts Dataset for Reproducibility-Oriented Sentiment Analysis
- National Institute on Aging PREPARE Challenge: Early Detection of Cognitive Impairment Using Speech -- The SpeechCARE Solution
- LLM-GROP: Visually Grounded Robot Task and Motion Planning with Large Language Models
- Cortex AISQL: A Production SQL Engine for Unstructured Data
- Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
- ZeroSim: Zero-Shot Analog Circuit Evaluation with Unified Transformer Embeddings
- On the Creativity of AI Agents
- Beyond human gold standards: A multimodel framework for automated abstract classification and information extraction
- Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
- AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
- CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
- Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
- Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
- Green AI: A systematic review and meta-analysis of its definitions, lifecycle models, hardware and measurement attempts
- CoLM: Collaborative Large Models via A Client-Server Paradigm
- Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
- Sampling and Loss Weights in Multi-Domain Training
- Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
- Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- TabRAG: Tabular Document Retrieval via Structured Language Representations
- PointCubeNet: 3D Part-level Reasoning with 3x3x3 Point Cloud Blocks
- Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
- Flexible Concept Bottleneck Model
- Can LLM Annotations Replace User Clicks for Learning to Rank?
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Who Is the Story About? Protagonist Entity Recognition in News
- Implicit Federated In-context Learning For Task-Specific LLM Fine-Tuning
- A Two-Stage System for Layout-Controlled Image Generation using Large Language Models and Diffusion Models
- MobileLLM-Pro Technical Report
- LLM For Loop Invariant Generation and Fixing: How Far Are We?
- CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
- How Well Do LLMs Understand Drug Mechanisms? A Knowledge + Reasoning Evaluation Dataset
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- ALIGN: A Vision-Language Framework for High-Accuracy Accident Location Inference through Geo-Spatial Neural Reasoning
- Meta-Learning-Driven GFlowNets for 3D Directional Modulation in Mobile Wireless Systems
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
- CSP4SDG: Constraint and Information-Theory Based Role Identification in Social Deduction Games with LLM-Enhanced Inference
- Scaling Laws and In-Context Learning: A Unified Theoretical Framework
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- LATTLE: LLM Attention Transplant for Transfer Learning of Tabular Data Across Disparate Domains
- Large Language Models Develop Novel Social Biases Through Adaptive Exploration
- Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
- L2T-Hyena: Enhancing State-Space Models with an Adaptive Learn-to-Teach Framework
- Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- DiagnoLLM: A Hybrid Bayesian Neural Language Framework for Interpretable Disease Diagnosis
- Adaptation and Fine-tuning with TabPFN for Travelling Salesman Problem
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- In-Context Learning Without Copying
- A Representation Sharpening Framework for Zero Shot Dense Retrieval
- The Peril of Preference: Why GRPO fails on Ordinal Rewards
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models
- 8bit-GPT: Exploring Human-AI Interaction on Obsolete Macintosh Operating Systems
- Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
- Generating Software Architecture Description from Source Code using Reverse Engineering and Large Language Model
- Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- Implementation of transformer-based LLMs with large-scale optoelectronic neurons on a CMOS image sensor platform
- What About Our Bug? A Study on the Responsiveness of NPM Package Maintainers
- Scientific judgment drifts over time in AI ideation
- Search Is Not Retrieval: Decoupling Semantic Matching from Contextual Assembly in RAG
- AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent
- REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment
- Software Defined Vehicle Code Generation: A Few-Shot Prompting Approach
- Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
- Cambrian-S: Towards Spatial Supersensing in Video
- Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
- Large language models replicate and predict human cooperation across experiments in game theory
- Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
- Differentially Private In-Context Learning with Nearest Neighbor Search
- Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
- An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
- Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
- TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
- LLM-enhanced Air Quality Monitoring Interface via Model Context Protocol
- Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis
- Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
- Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
- A systematic review of relation extraction task since the emergence of Transformers
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuning
- Two thousand years of the oracle problem. Insights from Ancient Delphi on the future of blockchain oracles
- Diffusion Language Models are Super Data Learners
- Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
- Automated Prompt Generation for Code Intelligence: An Empirical study and Experience in WeChat
- Control Barrier Function for Aligning Large Language Models
- Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
- Large Language Models as Information Sources: Distinctive Characteristics and Types of Low-Quality Information
- Analyzing the Power of Chain of Thought through Memorization Capabilities
- Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification
- QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
- Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
- The Curved Spacetime of Transformer Architectures
- ROBoto2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- Data-Efficient Adaptation and a Novel Evaluation Method for Aspect-based Sentiment Analysis
- Discrete Bayesian Sample Inference for Graph Generation
- Targeted Error Correction in Knowledge Distillation: Small Language Models Surpass GPT
- PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts
- Adaptive and Robust Data Poisoning Detection and Sanitization in Wearable IoT Systems using Large Language Models
- Modeling Hawkish-Dovish Latent Beliefs in Multi-Agent Debate-Based LLMs for Monetary Policy Decision Classification
- Next Token Knowledge Tracing: Exploiting Pretrained LLM Representations to Decode Student Behaviour
- Adapting General-Purpose Foundation Models for X-ray Ptychography in Low-Data Regimes
- Can Conversational AI Counsel for Change? A Theory-Driven Approach to Supporting Dietary Intentions in Ambivalent Individuals
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- LaRe: Latent Refocusing for Multimodal Reasoning
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
- IG-Pruning: Input-Guided Block Pruning for Large Language Models
- LLMs as Judges: Toward The Automatic Review of GSN-compliant Assurance Cases
- Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
- Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- ReleaseEval: A Benchmark for Evaluating Language Models in Automated Release Note Generation
- Can LLMs subtract numbers?
- The Collaboration Gap
- Memory-Efficient Training with In-Place FFT Implementation
- Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
- Routing-Based Continual Learning for Multimodal Large Language Models
- Random Initialization of Gated Sparse Adapters
- Context-Guided Decompilation: A Step Towards Re-executability
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- SemBench: A Benchmark for Semantic Query Processing Engines
- Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
- Exploring ChatGPT's Capabilities, Stability, Potential and Risks in Conducting Psychological Counseling through Simulations in School Counseling
- GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
- VayuChat: An LLM-Powered Conversational Interface for Air Quality Data Analytics
- On the Emergence of Induction Heads for In-Context Learning
- Aligning LLM agents with human learning and adjustment behavior: a dual agent approach
- Dynamic Multi-level Weighted Alignment Network for Zero-shot Sketch-based Image Retrieval
- G2rammar: Bilingual Grammar Modeling for Enhanced Text-attributed Graph Learning
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- Do Math Reasoning LLMs Help Predict the Impact of Public Transit Events?
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- A CPU-Centric Perspective on Agentic AI
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- Automated Invoice Data Extraction: Using LLM and OCR
- AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
- GDPR-Bench-Android: A Benchmark for Evaluating Automated GDPR Compliance Detection in Android
- EPARA: Parallelizing Categorized AI Inference in Edge Clouds
- Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
- Reasoning Planning for Language Models
- ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Reversal Invariance in Autoregressive Language Models
- Bayesian Network Structure Discovery Using Large Language Models
- TRISKELION-1: Unified Descriptive-Predictive-Generative AI
- LGCA: Enhancing Semantic Representation via Progressive Expansion
- ARC-GEN: A Mimetic Procedural Benchmark Generator for the Abstraction and Reasoning Corpus
- On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection
- From the Rock Floor to the Cloud: A Systematic Survey of State-of-the-Art NLP in Battery Life Cycle
- Rethinking Robust Adversarial Concept Erasure in Diffusion Models
- Traceable Drug Recommendation over Medical Knowledge Graphs
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- Synergistic Tensor and Pipeline Parallelism
- Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
- DRAMA: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries
- Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
- Language Modeling With Factorization Memory
- A Retrospect to Multi-prompt Learning across Vision and Language
- A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding
- Calibration Across Layers: Understanding Calibration Evolution in LLMs
- Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
- Detecting Data Contamination in LLMs via In-Context Learning
- Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations
- Masked Diffusion Captioning for Visual Feature Learning
- Pre-trained Forecasting Models: Strong Zero-Shot Feature Extractors for Time Series Classification
- SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
- An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
- Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
- Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling
- Hebrew Diacritics Restoration using Visual Representation
- LLMs as In-Context Meta-Learners for Model and Hyperparameter Selection
- Context Engineering 2.0: The Context of Context Engineering
- Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
- On the Role of Context for Discourse Relation Classification in Scientific Writing
- Towards Realistic Earth-Observation Constellation Scheduling: Benchmark and Methodology
- Unravelling the Mechanisms of Manipulating Numbers in Language Models
- Pragmatic Theories Enhance Understanding of Implied Meanings in LLMs
- Questionnaire meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
- RCScore: Quantifying Response Consistency in Large Language Models
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- Beyond Benchmarks: The Economics of AI Inference
- Learning Geometry: A Framework for Building Adaptive Manifold Models through Metric Optimization
- QuantumBench: A Benchmark for Quantum Problem Solving
- Predicate Renaming via Large Language Models
- Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
- SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: the Case of FrameNet Annotation
- Symbolically Scaffolded Play: Designing Role-Sensitive Prompts for Generative NPC Dialogue
- How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
- Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- Interpreting LLMs as Credit Risk Classifiers: Do Their Feature Explanations Align with Classical ML?
- EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- FARSIQA: Faithful and Advanced RAG System for Islamic Question Answering
- Standardization of Psychiatric Diagnoses -- Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System
- Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
- TextualVerifier: Verify TextGrad Step-by-Step
- Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
- MemSFT: Mitigating Alignment Tax with an External Parametric Memory
- A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries
- DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
- Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field
- LLM-Augmented Computational Phenotyping of Long Covid
- ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling
- Voice Memory for Agentic Speech Recognition
- LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
- HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
- Understanding Context Sampling in TabPFN on Small Tabular Datasets
- Scientific Knowledge Discovery in the Age of Large Language Models
- MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation
- The Innate Economic Preferences of Language Models
- GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
- When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
- Do What I Say: A Spoken Prompt Dataset for Instruction-Following
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- Training Language Models via Neural Cellular Automata
- Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
- Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
- Attention-Aligned Reasoning for Large Language Models
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- Generalization from Low- to Moderate-Resolution Spectra with Neural Networks for Stellar Parameter Estimation: A Case Study with DESI
- Symbol-Equivariant Recurrent Reasoning Models
- MedFusionT5: Cross-Modal Attention Boosts Semantic Quality and Reduces Hallucinations in Dental AI
- Convolutional neural networks in Vis–NIR chemometrics: From contradiction to conditional design
- HRM-Text: Efficient Pretraining Beyond Scaling
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- Evaluating Temporal Consistency in Multi-Turn Language Models
- Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Exploring the Intersection of AI, Language, and Law: A Bibliometric Analysis
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification
- SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
- France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- Generating units of cultural analysis with large language models: methods and validation for scalable cross-cultural research
- LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
- From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction
- On the Use of Large Language Models for Qualitative Synthesis
- How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
- CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
- Context Structure Reshapes the Representational Geometry of Language Models
- State of the Art of LLM-Enabled Interaction with Visualization
- Attentive multilayer fusion for vision transformers
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attribution
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- RL makes MLLMs see better than SFT
- Stay Tuned: Improving Sentiment Analysis and Stance Detection Using Large Language Models
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
- GReF: A Unified Generative Framework for Efficient Reranking via Ordered Multi-token Prediction
- Ideology-Based LLMs for Content Moderation
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA
- Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
- MemEIC: A Step Toward Continual and Compositional Knowledge Editing
- SeeingEye: Agentic Information Flow Unlocks Multimodal Reasoning In Text-only LLMs
- BioCoref: Benchmarking Biomedical Coreference Resolution with LLMs
- Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
- The use of LLMs to annotate data in management research: Foundational guidelines and warnings
- Future of AI Models: A Computational perspective on Model collapse
- MCP4IFC: IFC-Based Building Design Using Large Language Models
- Secure Retrieval-Augmented Generation against Poisoning Attacks
- StorageXTuner: An LLM Agent-Driven Automatic Tuning Framework for Heterogeneous Storage Systems
- Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
- What Really Matters in Matrix-Whitening Optimizers?
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- RiddleBench: A New Generative Reasoning Benchmark for LLMs
- Uniform Discrete Diffusion with Metric Path for Video Generation
- Bridging Tool Dependencies and Domain Knowledge: A Graph-Based Framework for In-Context Planning
- OrchDAG: Complex Tool Orchestration in Multi-Turn Interactions with Plan DAGs
- Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
- What Limits Agentic Systems Efficiency?
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
- BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Rethinking Visual Intelligence: Insights from Video Pretraining
- SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
- Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning
- LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
- MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation
- What do vision-language models see in the context? Investigating multimodal in-context learning
- Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?
- Evaluating LLMs on Generating Age-Appropriate Child-Like Conversations
- Enabling Near-realtime Remote Sensing via Satellite-Ground Collaboration of Large Vision-Language Models
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- Leveraging LLMs for Early Alzheimer's Prediction
- Auto prompting without training labels: An LLM cascade for product quality assessment in e-commerce catalogs
- Scalable GPU-Based Integrity Verification for Large Machine Learning Models
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- Information-Theoretic Discrete Diffusion
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Lifecycle-Aware code generation: Leveraging Software Engineering Phases in LLMs
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- Language Models for Longitudinal Clinical Prediction
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA
- Invoice Information Extraction: Methods and Performance Evaluation
- Artificial Intelligence and the Limits of Accumulation: Capital, Crisis, and the US Hegemonic Autumn in the World Market
- Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
- Evolution of the Tri-PDZ Domain in PSD95 (DLG-4 Gene)
- Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- A Survey on Efficient Vision-Language-Action Models
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?
- Unleashing Diverse Thinking Modes in LLMs through Multi-Agent Collaboration
- Evaluating Large Language Models for Stance Detection on Financial Targets from SEC Filing Reports and Earnings Call Transcripts
- Symbolic Neural Generation with Applications to Lead Discovery in Drug Design
- Large language model-based task planning for service robots: A review
- Provable test-time adaptivity and distributional robustness of in-context learning
- Multi-Stakeholder Alignment in LLM-Powered Collaborative AI Systems: A Multi-Agent Framework for Intelligent Tutoring
- A Survey on LLM Mid-Training
- LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
- TALM: Dynamic Tree-Structured Multi-Agent Framework with Long-Term Memory for Scalable Code Generation
- Can Language Models Compose Skills In-Context?
- Multi-Agent Conditional Diffusion Model with Mean Field Communication as Wireless Resource Allocation Planner
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies
- A Comprehensive Dataset for Human vs. AI Generated Text Detection
- Leveraging Large Language Models to Identify Conversation Threads in Collaborative Learning
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- SALSA: Single-pass Autoregressive LLM Structured Classification
- JiuTian Chuanliu: A Large Spatiotemporal Model for General-purpose Dynamic Urban Sensing
- A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
- Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
- Accelerating Materials Design via LLM-Guided Evolutionary Search
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
- A First Look at the Self-Admitted Technical Debt in Test Code: Taxonomy and Detection
- PortGPT: Towards Automated Backporting Using Large Language Models
- Harnessing the Power of Large Language Models for Software Testing Education: A Focus on ISTQB Syllabus
- From Slides to Chatbots: Enhancing Large Language Models with University Course Materials
- You Don't Need Prompt Engineering Anymore: The Prompting Inversion
- LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test Generation
- Pruning and Quantization Impact on Graph Neural Networks
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging
- Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
- CMOMgen: Complex Multi-Ontology Alignment via Pattern-Guided In-Context Learning
- LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- RETuning: Upgrading Inference-Time Scaling for Stock Movement Prediction with Large Language Models
- Head Pursuit: Probing Attention Specialization in Multimodal Transformers
- SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
- MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection
- Unified token representations for sequential decision models
- Large Language Models as Model Organisms for Human Associative Learning
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Compressing Many-Shots in In-Context Learning
- Securing AI Agent Execution
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Designing and Evaluating Hint Generation Systems for Science Education
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- The Virtues of Brevity: Avoid Overthinking in Parallel Test-Time Reasoning
- Dynamic Retriever for In-Context Knowledge Editing via Policy Optimization
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
- A Scalable, Causal, and Energy Efficient Framework for Neural Decoding with Spiking Neural Networks
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- ARC-Encoder: learning compressed text representations for large language models
- Large Language Models for Fault Localization: An Empirical Study
- Hardware-Aware DNN Compression for Homogeneous Edge Devices
- FreeChunker: A Cross-Granularity Chunking Framework
- Exponential Convergence Guarantees for Iterative Markovian Fitting
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models
- Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
- LM-mixup: Text Data Augmentation via Language Model based Mixup
- MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure Elucidation
- Using Large Language Models for Abstraction of Planning Domains - Extended Version
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Leveraging the Power of Large Language Models in Entity Linking via Adaptive Routing and Targeted Reasoning
- Meta-Learning for Cross-Task Generalization in Protein Mutation Property Prediction
- The Impact of Negated Text on Hallucination with Large Language Models
- On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
- HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
- Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
- NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew
- Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
- Latent Space Factorization in LoRA
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- COLA: Continual Learning via Autoencoder Retrieval of Adapters
- CPSVD: Enhancing Large Language Model Compression via Column-Preserving Singular Value Decomposition
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
- KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Selecting and Combining Large Language Models for Scalable Code Clone Detection
- Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
- News-Aware Direct Reinforcement Trading for Financial Markets
- When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
- SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations
- VeFA: Vector-Based Feature Space Adaptation for Robust Model Fine-Tuning
- Transformers are almost optimal metalearners for linear classification
- ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
- LAPRAD: LLM-Assisted PRotocol Attack Discovery
- RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
- When Your AI Agent Succumbs to Peer-Pressure: Studying Opinion-Change Dynamics of LLMs
- E-Test: E'er-Improving Test Suites
- The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
- XGen-Q: An Explainable Domain-Adaptive LLM Framework with Retrieval-Augmented Generation for Software Security
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting
- Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
- Fetch.ai: An Architecture for Modern Multi-Agent Systems
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
- Large language models for folktale type automation based on motifs: Cinderella case study
- VAPU: System for Autonomous Legacy Code Modernization
- Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization
- Simple and Efficient Heterogeneous Temporal Graph Neural Network
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
- From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering
- Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
- OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- ActivationReasoning: Logical Reasoning in Latent Activation Spaces
- Model Context Contracts - MCP-Enabled Framework to Integrate LLMs With Blockchain Smart Contracts
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
- DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
- Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
- Automatic Prompt Generation via Adaptive Selection of Prompting Techniques
- Rethinking PCA Through Duality
- Online In-Context Distillation for Low-Resource Vision Language Models
- Planned Diffusion
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- MEG-GPT: A transformer-based foundation model for magnetoencephalography data
- CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows
- Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
- DynaQuery: A Self-Adapting Framework for Querying Structured and Multimodal Data
- Glyph: Scaling Context Windows via Visual-Text Compression
- Closing the Sim2Real Performance Gap in RL
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
- Layer Specialization Underlying Compositional Reasoning in Transformers
- Multilingual Clinical NER for Diseases and Medications Recognition in Cardiology Texts using BERT Embeddings
- Diffusion Models as Dataset Distillation Priors
- 3S-Trader: A Multi-LLM Framework for Adaptive Stock Scoring, Strategy, and Selection in Portfolio Optimization
- The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
- Strengthening LLMs for Tabular Prediction with Structural Priors
- RubiSCoT: A Framework for AI-Supported Academic Assessment
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
- Variance-Reduction Guidance: Sampling Trajectory Optimization for Diffusion Models
- Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
- An Evaluation of LLMs Inference on Popular Single-board Computers
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
- Soft-Masked Diffusion Language Models
- DeTAILS: Deep Thematic Analysis with Iterative LLM Support
- PEACE: Towards Efficient Project-Level Efficiency Optimization via Hybrid Code Editing
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Saber: An Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model
- I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- Reasoning Distillation and Structural Alignment for Improved Code Generation
- Annotation-Efficient Universal Honesty Alignment
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
- A Systematic Literature Review of the Use of GenAI Assistants for Code Comprehension: Implications for Computing Education Research and Practice
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
- Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding
- Learning to play: A Multimodal Agent for 3D Game-Play
- Efficient High-Accuracy PDEs Solver with the Linear Attention Neural Operator
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
- Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
- An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review
- Prompt Optimization via Retrieved Reasoning Assets and Multi-Agent Analysis
- Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
- Safe and Efficient In-Context Learning via Risk Control
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
- Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
- The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
- KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
- Policy Transfer for Continuous-Time Reinforcement Learning: A (Rough) Differential Equation Approach
- Latent Topic Synthesis: Leveraging LLMs for Electoral Ad Analysis
- PrivacyPAD: A Reinforcement Learning Framework for Dynamic Privacy-Aware Delegation
- An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
- DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- Terra: Explorable Native 3D World Model with Point Latents
- ChangingGrounding: 3D Visual Grounding in Changing Scenes
- Predicting Task Performance with Context-aware Scaling Laws
- A Novel GPT-Based Framework for Anomaly Detection in System Logs
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- Cognitive-Aligned Spatio-Temporal Large Language Models For Next Point-of-Interest Prediction
- ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
- Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
- State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
- Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Your Next Token Prediction: A Multilingual Benchmark for Personalized Response Generation
- Suicidal Comment Tree Dataset: Enhancing Risk Assessment and Prediction Through Contextual Analysis
- Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
- PluriHop: Exhaustive, Recall-Sensitive QA over Distractor-Rich Corpora
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- LLM-ERM: Sample-Efficient Program Learning via LLM-Guided Search
- Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Spatial Computing Communications for Multi-User Virtual Reality in Distributed Mobile Edge Computing Network
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
- From Attention to Disaggregation: Tracing the Evolution of LLM Inference
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
- Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models
- Inferred global dense residue transition graphs from primary structure sequences enable protein interaction prediction via directed graph convolutional neural networks
- Multimodal Function Vectors for Visual Relations
- David vs. Goliath: A comparative study of different-sized LLMs for code generation in the domain of automotive scenario generation
- FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
- Quantifying Phonosemantic Iconicity Distributionally in 6 Languages
- Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping Games
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- Scaling Vision Transformers for Functional MRI with Flat Maps
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Big Reasoning with Small Models: Instruction Retrieval at Inference Time
- Rethinking Evaluation in the Era of Time Series Foundation Models: (Un)known Information Leakage Challenges
- The Role of Computing Resources in Publishing Foundation Model Research
- MemoTime: Memory-Augmented Temporal Knowledge Graph Enhanced Large Language Model Reasoning
- Element2Vec: Build Chemical Element Representation from Text for Property Prediction
- LLM one-shot style transfer for Authorship Attribution and Verification
- F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
- FACTS: Table Summarization via Offline Template Generation with Agentic Workflows
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Retrieval-in-the-Chain: Bootstrapping Large Language Models for Generative Retrieval
- NeuroRVQ: Multi-Scale EEG Tokenization for Generative Large Brainwave Models
- Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan's Intelligent Interaction Systems
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
- Convergence, design and training of continuous-time dropout as a random batch method
- Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- Universal Image Restoration Pre-training via Masked Degradation Classification
- TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- A Matter of Representation: Towards Graph-Based Abstract Code Generation
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- Text Anomaly Detection with Simplified Isolation Kernel
- D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
- RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
- Schema for In-Context Learning
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
- Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
- COSTAR-A: A prompting framework for enhancing Large Language Model performance on Point-of-View questions
- Artificial Intelligence Virtual Cells: From Measurements to Decisions across Modality, Scale, Dynamics, and Evaluation
- When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
- Cautious Weight Decay
- IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion
- Improving Generative Behavior Cloning via Self-Guidance and Adaptive Chunking
- A Hierarchical Quantized Tokenization Framework for Task-Adaptive Graph Representation Learning
- Traveling Salesman-Based Token Ordering Improves Stability in Homomorphically Encrypted Language Models
- PromptFlow: Training Prompts Like Neural Networks
- Encapsulating Textual Contents into a MOC data Structure for Advanced Applications
- Reinforced Preference Optimization for Recommendation
- Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
- A Survey on Parallel Reasoning
- Single chip 1 Tb/s optical transmitter with inverse designed input and output couplers
- Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time Adaptation
- Credal Transformer: A Principled Approach for Quantifying and Mitigating Hallucinations in Large Language Models
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- Stratos: An End-to-End Distillation Pipeline for Customized LLMs under Distributed Cloud Environments
- HiCoTraj:Zero-Shot Demographic Reasoning via Hierarchical Chain-of-Thought Prompting from Trajectory
- Empowering LLM Agents with Geospatial Awareness: Toward Grounded Reasoning for Wildfire Response
- Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
- MedKGEval: A Knowledge Graph-Based Multi-Turn Evaluation Framework for Open-Ended Patient Interactions with Clinical LLMs
- SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
- Leveraging Language Semantics for Collaborative Filtering with TextGCN and TextGCN-MLP: Zero-Shot vs In-Domain Performance
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Deep Research Brings Deeper Harm
- MeTA-LoRA: Data-Efficient Multi-Task Fine-Tuning for Large Language Models
- AgentCaster: Reasoning-Guided Tornado Forecasting
- Exploring Artificial Intelligence and Culture: Methodology for a comparative study of AI's impact on norms, trust, and problem-solving across academic and business environments
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- Investigating Large Language Models' Linguistic Abilities for Text Preprocessing
- VeriCite: Towards Reliable Citations in Retrieval-Augmented Generation via Rigorous Verification
- Unlocking the Potential of Diffusion Language Models through Template Infilling
- Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- Vision-LLMs for Spatiotemporal Traffic Forecasting
- A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
- Is Implicit Knowledge Enough for LLMs? A RAG Approach for Tree-based Structures
- Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States
- LogiNumSynth: Synthesizing Joint Logical-Numerical Reasoning Problems for Language Models
- J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
- ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- Bolster Hallucination Detection via Prompt-Guided Data Augmentation
- QLENS: Towards A Quantum Perspective of Language Transformers
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- A Two-Step, Multidimensional Account of Deception in Language Models
- Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
- Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
- PAGE: Prompt Augmentation for text Generation Enhancement
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- Cognitive Load Traces as Symbolic and Visual Accounts of Deep Model Cognition
- Direct Multi-Token Decoding
- The Command Line GUIde: Graphical Interfaces from Man Pages via AI
- Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
- RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- High-Fidelity Speech Enhancement via Discrete Audio Tokens
- Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment Planning
- Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems
- HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication
- Merlin's Whisper: Enabling Efficient Reasoning in LLMs via Black-box Adversarial Prompting
- Assessing Large Language Models for Structured Medical Order Extraction
- ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
- Harnessing Consistency for Robust Test-Time LLM Ensemble
- Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- Softmax ≥ Linear: Transformers may learn to classify in-context by kernel gradient descent
- Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
- Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
- DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
- PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
- CauchyNet: Compact and Data-Efficient Learning using Holomorphic Activation Functions
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- A-IPO: Adaptive Intent-driven Preference Optimization
- Serialized EHR make for good text representations
- Taking a SEAT: Predicting Value Interpretations from Sentiment, Emotion, Argument, and Topic Annotations
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold
- GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
- Are LLMs Better GNN Helpers? Rethinking Robust Graph Learning under Deficiencies with Iterative Refinement
- An agentic artificially intelligent X-ray scientist
- Exploring Large Language Models for Financial Applications: Techniques, Performance, and Challenges with FinMA
- Doc-to-Atom: Learning to Compile and Compose Memory Atoms
- Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
- Accelerating Attention with Basis Decomposition
- Small edits, large models: How Wikipedia advocacy shapes LLM values
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- SeeingSounds: Learning Audio-to-Visual Alignment via Text
- Classifier-Augmented Generation for Structured Workflow Prediction
- An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
- Stability of Transformers under Layer Normalization
- Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
- PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection
- Patentformer: A demonstration of AI-assisted automated patent drafting
- Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
- Doc2Query++: Topic-Coverage based Document Expansion and its Application to Dense Retrieval via Dual-Index Fusion
- Titans Revisited: A Lightweight Reimplementation and Critical Analysis of a Test-Time Memory Model
- Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
- StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- AdaPM: a Partial Momentum Algorithm for LLM Training
- Provable Watermarking for Data Poisoning Attacks
- Efficient Resource-Constrained Training of Vision Transformers via Subspace Optimization
- LitE-SQL: A Lightweight and Efficient Text-to-SQL Framework with Vector-based Schema Linking and Execution-Guided Self-Correction
- FrameEOL: Semantic Frame Induction using Causal Language Models
- ICL-Router: In-Context Learned Model Representations for LLM Routing
- MEC3O: Multi-Expert Consensus for Code Time Complexity Prediction
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- Repairing Regex Vulnerabilities via Localization-Guided Instructions
- Promptimizer: User-Led Prompt Optimization for Personal Content Classification
- MASA: LLM-Driven Multi-Agent Systems for Autoformalization
- Creation of the Chinese Adaptive Policy Communication Corpus
- Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
- The Idola Tribus of AI: Large Language Models tend to perceive order where none exists
- Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
- Modeling Layered Consciousness with Multi-Agent Large Language Models
- Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
- Efficient Autoregressive Inference for Transformer Probabilistic Models
- Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
- Score-Based Density Estimation from Pairwise Comparisons
- Understanding the Effects of Domain Finetuning on LLMs
- Task-Level Insights from Eigenvalues across Sequence Models
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- NL2GenSym: Natural Language to Generative Symbolic Rules for SOAR Cognitive Architecture via Large Language Models
- Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
- LLM Based Long Code Translation using Identifier Replacement
- Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks
- PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document Retrieval
- TinyGraphEstimator: Adapting Lightweight Language Models for Graph Structure Inference
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Towards Neurocognitive-Inspired Intelligence: From AI's Structural Mimicry to Human-Like Functional Cognition
- Graph Diffusion Transformers are In-Context Molecular Designers
- Scaling Laws for Code: A More Data-Hungry Regime
- Opponent Shaping in LLM Agents
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- LLM-Assisted Web Measurements
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
- Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention
- Wavefunction Flows: Efficient Quantum Simulation of Continuous Flow Models
- Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
- Position: Privacy Is Not Just Memorization!
- On the Relationship Between the Choice of Representation and In-Context Learning
- Improving Reasoning for Diffusion Language Models via Group Diffusion Policy Optimization
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- MeSH: Memory-as-State-Highways for Recursive Transformers
- AppForge: From Assistant to Independent Developer -- Are GPTs Ready for Software Development?
- Drift No More? Context Equilibria in Multi-Turn LLM Interactions
- Prompts Generalize with Low Data: Non-vacuous Generalization Bounds for Optimizing Prompts with More Informative Priors
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Models
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
- DACIP-RC: Domain Adaptive Continual Instruction Pre-Training via Reading Comprehension on Business Conversations
- TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
- In-Context Clustering with Large Language Models
- Training-Free Group Relative Policy Optimization
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- Toward Reliable Clinical Coding with Language Models: Verification and Lightweight Adaptation
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning
- LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning
- Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
- Benchmarking is Broken -- Don't Let AI be its Own Judge
- MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
- Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
- Accelerating Inference for Multilayer Neural Networks with Quantum Computers
- CARPAS: Towards Content-Aware Refinement of Provided Aspects for Summarization in Large Language Models
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
- Prompt Optimization Across Multiple Agents for Representing Diverse Human Populations
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- Textual interpretation of transient image classifications from large language models
- When Machines Meet Each Other: Network Effects and the Strategic Role of History in Multi-Agent AI
- Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
- AMAS: Adaptively Determining Communication Topology for LLM-based Multi-Agent System
- End-to-End Test-Time Training for Long Context
- BOTANIC-0: a series of foundation models for plant genomic data
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Mid-Training of Large Language Models: A Survey
- Dual Goal Representations
- Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
- Heptapod: Language Modeling on Visual Signals
- Aligning Large Language Models via Fully Self-Synthetic Data
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Fine-Grained Emotion Recognition via In-Context Learning
- InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
- Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
- AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
- Incoherence in Goal-Conditioned Autoregressive Models
- Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
- Test-Time Scaling of Reasoning Models for Machine Translation
- Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
- Training Dynamics Impact Post-Training Quantization Robustness
- SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
- TabPFN-Wide: Continued Pre-Training for Extreme Feature Counts
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Taxonomy of User Needs and Actions
- Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
- Prompt reinforcing for long-term planning of large language models
- MaNGO - Adaptable Graph Network Simulators via Meta-Learning
- DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
- Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling
- Mixture of Neuron Experts
- Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
- Syn-Diag: An LLM-based Synergistic Framework for Generalizable Few-shot Fault Diagnosis on the Edge
- Membership Inference Attacks on Tokenizers of Large Language Models
- YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
- vAttention: Verified Sparse Attention
- AutoPentester: An LLM Agent-based Framework for Automated Pentesting
- BuilderBench -- A benchmark for generalist agents
- Sci-Phi: A Large Language Model Spatial Audio Descriptor
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- Mission Impossible: Feedback-Guided Dynamic Interactive Planning for Improving Reasoning on LLMs
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- Information-Theoretic Policy Pre-Training with Empowerment
- MASA: Rethinking the Representational Bottleneck in LoRA with Multi-A Shared Adaptation
- The New Quant: A Survey of Large Language Models in Financial Prediction and Trading
- Prototype-Based Dynamic Steering for Large Language Models
- Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
- Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
- SocialNLI: A Dialogue-Centric Social Inference Dataset
- InvThink: Premortem Reasoning for Safer Language Models
- Efficient Prediction of Pass@k Scaling in Large Language Models
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- Adjusting the Output of Decision Transformer with Action Gradient
- Finish First, Perfect Later: Test-Time Token-Level Cross-Validation for Diffusion Large Language Models
- Staircase Streaming for Low-Latency Multi-Agent Inference
- KEEP: Integrating Medical Ontologies with Clinical Data for Robust Code Embeddings
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
- LMM-Incentive: Large Multimodal Model-based Incentive Design for User-Generated Content in Web 3.0
- LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows
- ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRA
- Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- GILT: An LLM-Free, Tuning-Free Graph Foundational Model for In-Context Learning
- Conditional Representation Learning for Customized Tasks
- ContextNav: Towards Agentic Multimodal In-Context Learning
- Learning Linear Regression with Low-Rank Tasks in-Context
- Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
- Multi-Agent Collaborative Intelligence: Dual-Dial Control for Reliable LLM Reasoning
- Forking-Sequences
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
- Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Pulp Motion: Framing-aware multimodal camera and human motion generation
- Making Mathematical Reasoning Adaptive
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
- Can LLMs Detect Ambiguous Plural Reference? An Analysis of Split-Antecedent and Mereological Reference
- Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
- Kernel ridge regression under power-law data: spectrum and generalization
- COSMIR: Chain Orchestrated Structured Memory for Iterative Reasoning over Long Context
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- Evaluating Self-Supervised Speech Models via Text-Based LLMS
- ActiveMark: on watermarking of visual foundation models via massive activations
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- LLM Based Bayesian Optimization for Prompt Search
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- BLADE: Bias-Linked Adaptive DEbiasing
- Proof Artifact Co-training for Theorem Proving with Language Models
- Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- Evaluation of Clinical Trials Reporting Quality using Large Language Models
- Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
- Thai Semantic End-of-Turn Detection for Real-Time Voice Agents
- Small Language Models for Emergency Departments Decision Support: A Benchmark Study
- Critical appraisal of artificial intelligence for rare-event recognition: principles and pharmacovigilance case studies
- LongTail-Swap: benchmarking language models' abilities on rare words
- Large Language Models Hallucination: A Comprehensive Survey
- Exact Causal Attention with 10% Fewer Operations
- Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness
- Multi Language Models for On-the-Fly Syntax Highlighting
- Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- LLM as an Algorithmist: Enhancing Anomaly Detectors via Programmatic Synthesis
- Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
- PoseGaze-AHP: A Knowledge-Based 3D Dataset for AI-Driven Ocular and Postural Diagnosis
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
- Mechanistic Interpretability of Socio-Political Frames in Language Models
- Allocation of Parameters in Transformers
- Prompt Balance Matters: Understanding How Imbalanced Few-Shot Learning Affects Multilingual Sense Disambiguation in LLMs
- Evolutionary Computation as Natural Generative AI
- On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- Generating High-Level Test Cases from Requirements using LLM: An Industry Study
- REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
- Self-Speculative Masked Diffusions
- Efficient Test-Time Scaling for Small Vision-Language Models
- A Qualitative Comparative Evaluation of Cognitive and Generative Theories
- Know Thyself? On the Incapability and Implications of AI Self-Recognition
- Neural Correlates of Language Models Are Specific to Human Language
- Evaluating Embedding Frameworks for Scientific Domain
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
- IndiCASA: A Dataset and Bias Evaluation Framework in LLMs Using Contrastive Embedding Similarity in the Indian Context
- TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
- Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- Automated Constraint Specification for Job Scheduling by Regulating Generative Model with Domain-Specific Representation
- HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Less LLM, More Documents: Searching for Improved RAG
- Geolog-IA: Conversational System for Academic Theses
- SoT: Structured-of-Thought Prompting Guides Multilingual Reasoning in Large Language Models
- Deep Generative Continual Learning using Functional LoRA: FunLoRA
- MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
- ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- A Simple but Effective Elaborative Query Reformulation Approach for Natural Language Recommendation
- A Study of Rule Omission in Raven's Progressive Matrices
- Homophily-induced Emergence of Biased Structures in LLM-based Multi-Agent AI Systems
- Automatic Building Code Review: A Case Study
- Fine-tuning LLMs with variational Bayesian last layer for high-dimensional Bayesian optimization
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
- GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
- Backdoor Attacks Against Speech Language Models
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- Integrating AI and Ensemble Forecasting: Explainable Materials Planning with Scorecards and Trend Insights for a Large-Scale Manufacturer
- Visual Self-Refinement for Autoregressive Models
- Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
- The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
- Conversational Agents for Building Energy Efficiency -- Advising Housing Cooperatives in Stockholm on Reducing Energy Consumption
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
- Hybrid Training for Vision-Language-Action Models
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- Memory-Augmented Log Analysis with Phi-4-mini: Enhancing Threat Detection in Structured Security Logs
- Structuring Reasoning for Complex Rules Beyond Flat Representations
- TokMem: Tokenized Procedural Memory for Large Language Models
- Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
- Generative AI for subgrid turbulence in large-eddy simulations
- Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
- Neu-RadBERT for Enhanced Diagnosis of Brain Injuries and Conditions
- How Foundational are Foundation Models for Time Series Forecasting?
- Reasoning-Aware Prompt Orchestration: A Foundation Model for Multi-Agent Language Model Coordination
- ICL Optimized Fragility
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
- RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
- TVS Sidekick: Challenges and Practical Insights from Deploying Large Language Models in the Enterprise
- ACT: Agentic Classification Tree
- Communication-Efficient and Accurate Approach for Aggregation in Federated Low-Rank Adaptation
- Are neural scaling laws leading quantum chemistry astray?
- MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
- Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
- Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models
- Nephrobase Cell+: Multimodal Single-Cell Foundation Model for Decoding Kidney Biology
- LLM Agents for Knowledge Discovery in Atomic Layer Processing
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- 90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development
- CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models
- End-to-End Aspect-Guided Review Summarization at Scale
- Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
- Indirect Attention: Turning Context Misalignment into a Feature
- Using GPT to build a Project Management assistant for Jira environments
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
- Better Privilege Separation for Agents by Restricting Data Types
- Accelerating LLM Inference with Precomputed Query Storage
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Revoking Amnesia: RL-based Trajectory Optimization to Resurrect Erased Concepts in Diffusion Models
- RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
- Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- Galton's Law of Mediocrity: Why Large Language Models Regress to the Mean and Fail at Creativity in Advertising
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- Test time training enhances in-context learning of nonlinear functions
- Transformer-Based Neural Networks Backflow for Strongly Correlated Electronic Structure
- The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
- QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
- Transformers through the lens of support-preserving maps between measures
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Explainable Fault Localization for Programming Assignments via LLM-Guided Annotation
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
- Submodular Context Partitioning and Compression for In-Context Learning
- Directed Information γ-covering: An Information-Theoretic Framework for Context Engineering
- Hierarchical Reasoning Models: Perspectives and Misconceptions
- LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
- Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
- Uncovering Zero-Shot Generalization Gaps in Time-Series Foundation Models Using Real-World Videos
- SafePassage: High-Fidelity Information Extraction with Black Box LLMs
- Judging by Appearances? Auditing and Intervening Vision-Language Models for Bail Prediction
- MetaChest: Generalized few-shot learning of pathologies from chest X-rays
- Information Design With Large Language Models
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling View
- Towards Structured Knowledge: Advancing Triple Extraction from Regional Trade Agreements using Large Language Models
- A Cartography of Open Collaboration in Open Source AI: Mapping Practices, Motivations, and Governance in 14 Open Large Language Model Projects
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks
- Towards Reliable Generation of Executable Workflows by Foundation Models
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
- GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
- CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
- Maximizing Parallelism in Distributed Training for Huge Neural Networks
- SecInfer: Preventing Prompt Injection via Inference-time Scaling
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- How Well Do LLMs Imitate Human Writing Style?
- Inductive Bias and Spectral Properties of Single-Head Attention in High Dimensions
- Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
- Environment-Aware Satellite Image Generation with Diffusion Models
- Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
- HAPT: Heterogeneity-Aware Automated Parallel Training on Heterogeneous Clusters
- Pushing LLMs to Their Logical Reasoning Bound: The Role of Data Reasoning Intensity
- KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical Reasoning
- FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits
- Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
- PRIVMARK: Private Large Language Models Watermarking with MPC
- Prompting Robot Teams with Natural Language
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
- DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability
- Agentic Services Computing
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
- Comparing Open-Source and Commercial LLMs for Domain-Specific Analysis and Reporting: Software Engineering Challenges and Design Trade-offs
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Training Dynamics of Parametric and In-Context Knowledge Utilization in Language Models
- PEARL: Performance-Enhanced Aggregated Representation Learning
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
- VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference
- Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm
- Conda: Column-Normalized Adam for Training Large Language Models Faster
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
- Watermarking Diffusion Language Models
- InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- NeMo: Needle in a Montage for Video-Language Understanding
- Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
- Act as an expert in psychometry. The evaluation of large language models utility in psychological tests cross-cultural adaptations
- Muon: Training and Trade-offs with Latent Attention and MoE
- Personalized Vision via Visual In-Context Learning
- Rethinking and Benchmarking Large Language Models for Graph Reasoning
- LLM-Assisted News Discovery in High-Volume Information Streams: A Case Study
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- Free Access to World News: Reconstructing Full-Text Articles from GDELT
- Reinforcement Mid-Training
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- Does Weak-to-strong Generalization Happen under Spurious Correlations?
- Detecting and Rectifying Noisy Labels: A Similarity-based Approach
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
- Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
- From Neural Networks to Logical Theories: The Correspondence between Fibring Modal Logics and Fibring Neural Networks
- Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
- Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
- Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
- Anchored Supervised Fine-Tuning
- Time-Shifted Token Scheduling for Symbolic Music Generation
- LocoFormer: Generalist Locomotion via Long-context Adaptation
- A Weather Foundation Model for the Power Grid
- AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
- StrucADT: Generating Structure-controlled 3D Point Clouds with Adjacency Diffusion Transformer
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- Training Optimal Large Diffusion Language Models
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- A Computational Perspective on NeuroAI and Synthetic Biological Intelligence
- MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models
- MemMamba: Rethinking Memory Patterns in State Space Model
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- Privy: Envisioning and Mitigating Privacy Risks for Consumer-facing AI Product Concepts
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
- The Impact of Role Design in In-Context Learning for Large Language Models
- Democratizing AI scientists using ToolUniverse
- MimiTalk: Revolutionizing Qualitative Research with Dual-Agent AI
- Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
- A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- Towards Monotonic Improvement in In-Context Reinforcement Learning
- From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs
- PDE-Transformer: A Continuous Dynamical Systems Approach to Sequence Modeling
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- Limit Analysis for Symbolic Multi-step Reasoning Tasks with Information Propagation Rules Based on Transformers
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
- Effective Quantization of Muon Optimizer States
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
- Planning with Unified Multimodal Models
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning
- BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
- Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
- Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2
- Towards Human-interpretable Explanation in Code Clone Detection using LLM-based Post Hoc Explainer
- Ringleader ASGD: The First Asynchronous SGD with Optimal Time Complexity under Data Heterogeneity
- Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention
- Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
- What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
- Ethische Betrachtungen der automatisierten Textanalyse <b>: Fortschritte, Risiken und die Notwendigkeit eines Gleichgewichts</b>
- Scale-Wise VAR is Secretly Discrete Diffusion
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Category Discovery: An Open-World Perspective
- Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks
- Partial Parameter Updates for Efficient Distributed Training
- A model of errors in transformers
- Meta-Awareness Enhances Reasoning Models: Self-Alignment Reinforcement Learning
- Stochastic activations
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- Context Parametrization with Compositional Adapters
- Multi-Agent Path Finding via Offline RL and LLM Collaboration
- FoodSEM: Large Language Model Specialized in Food Named-Entity Linking
- Multilingual Vision-Language Models, A Survey
- Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
- The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
- Teaching Transformers to Solve Combinatorial Problems through Efficient Trial & Error
- Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
- A High-Capacity and Secure Disambiguation Algorithm for Neural Linguistic Steganography
- A Large-Scale Dataset and Citation Intent Classification in Turkish with LLMs
- Enhancing Low-Rank Adaptation with Structured Nonlinear Transformations
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineering
- Can LLMs Solve and Generate Linguistic Olympiad Puzzles?
- Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies
- UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
- AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
- In-Context Learning can Perform Continual Learning Like Humans
- Why Chain of Thought Fails in Clinical Text Understanding
- Semantic-Inductive Attribute Selection for Zero-Shot Learning
- Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
- RLP: Reinforcement as a Pretraining Objective
- Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
- MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- Towards Transparent AI: A Survey on Explainable Language Models
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Plan2Evolve: LLM Self-Evolution for Improved Planning Capability via Automated Domain Generation
- A circuit for predicting hierarchical structure in-context in Large Language Models
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- On Code-Induced Reasoning in LLMs
- GraphPFN: A Prior-Data Fitted Graph Foundation Model
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
- Foundation models for high-energy physics
- Data-Centric Elastic Pipeline Parallelism for Efficient Long-Context LLM Training
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- Explaining Fine Tuned LLMs via Counterfactuals A Knowledge Graph Driven Framework
- PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- Disagreements in Reasoning: How a Model's Thinking Process Dictates Persuasion in Multi-Agent Systems
- Predicting LLM Reasoning Performance with Small Proxy Model
- A short survey on almost orthogonal vectors in a few specific large dimensions
- GALAX: Graph-Augmented Language Model for Explainable Reinforcement-Guided Subgraph Reasoning in Precision Medicine
- Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search
- Distilling Many-Shot In-Context Learning into a Cheat Sheet
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching
- Towards Atoms of Large Language Models
- Concise and Sufficient Sub-Sentence Citations for Retrieval-Augmented Generation
- Generative AI for FFRDCs
- LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
- It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
- RJE: A Retrieval-Judgment-Exploration Framework for Efficient Knowledge Graph Question Answering with LLMs
- On Theoretical Interpretations of Concept-Based In-Context Learning
- Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos
- LAVA: Explainability for Unsupervised Latent Embeddings
- When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
- CHARM: Control-point-based 3D Anime Hairstyle Auto-Regressive Modeling
- An LLM-based Agentic Framework for Accessible Network Control
- MARS: toward more efficient multi-agent collaboration for LLM reasoning
- PromptDebt: A Comprehensive Study of Technical Debt Across LLM Projects
- Document Summarization with Conformal Importance Guarantees
- DRES: Benchmarking LLMs for Disfluency Removal
- Thinking Augmented Pre-training
- Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder
- GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
- AMLA: MUL by ADD in FlashAttention Rescaling
- Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- DAOpt: Modeling and Evaluation of Data-Driven Optimization under Uncertainty with LLMs
- BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
- Polarity Detection of Sustainable Detection Goals in News Text
- Large Language Models for Real-World IoT Device Identification
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- Linear Transformers Implicitly Discover Unified Numerical Algorithms
- Thinking While Listening: Simple Test Time Scaling For Audio Classification
- RoboSSM: Scalable In-context Imitation Learning via State-Space Models
- Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- DELM: a Python toolkit for Data Extraction with Language Models
- Mamba Modulation: On the Length Generalization of Mamba
- GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- Nano Bio-Agents (NBA): Small Language Model Agents for Genomics
- Uncertainty in Semantic Language Modeling with PIXELS
- Confidence Calibration in Large Language Model-Based Entity Matching
- Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
- LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
- Training Skills Like Parameters via Self-Supervised Semantic Diffusion
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
- FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
- What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
- Revisiting Predictive Process Monitoring in the Age of Foundation Models: A Comparative Study of Sequence, Tabular, and LLM Approaches
- GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
- Towards joint scaling laws with optimal batch size schedules
- Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
- A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
- A foundation model of numerical intelligence with cross-disciplinary generalization
- Can Vision-Language Models Reason about AI Edits in Images?
- Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- Explaining Data Mixing Scaling Laws
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
- Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
- mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- Persuading large language models to comply with objectionable requests
- Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation
- RELISH: LLM REgression with a Latent Iterative State Head
- Verbalizing LLM's Higher-order Uncertainty via Imprecise Probabilities
- A Theory of Appropriateness That Accounts for Norms of Rationality
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- Benchmark for Assessing Olfactory Perception of Large Language Models
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- Non-Parametric Structural Priors for Geometry Theorem Prediction
- Transporting Task Vectors across Different Architectures without Training
- Use What You Know: Causal Foundation Models with Partial Graphs
- A large-scale evaluation of commonsense knowledge in humans and large language models
- epiGPTope: A Machine Learning-Based Epitope Generator and Classifier
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- Beyond computational equivalence: the behavioral inference principle for machine consciousness
- Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
- A Retail-Corpus for Aspect-Based Sentiment Analysis with Large Language Models
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Language Models Coupled with Metacognition Can Outperform Reasoning Models
- SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models
- A unified large language model–based framework for heterogeneous PV image diagnosis
- AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation
- Riemannian Optimization for LoRA on the Stiefel Manifold
- Generative AI for Born-Digital Collections: A Case Study of Metadata Management in an Institutional Repository
- Training Neural Networks with Fixed Sparse Masks
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
- VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
- Helixer: ab initio prediction of primary eukaryotic gene models combining deep learning and a hidden Markov model
- Limitations of Normalization in Attention Mechanism
- A multimodal GeoAI approach to combining text with spatiotemporal features for enhanced relevance classification of social media posts in disaster response
- Cascaded Text Generation with Markov Transformers
- When Does Preconditioning Help or Hurt Generalization?
- Advancing Thesaurus Construction With Generative AI: A Structured Approach to Synonym Identification and Scope Note Development
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
- Reinforcement Learning on Pre-Training Data
- A Knowledge Graph and a Tripartite Evaluation Framework Make Retrieval-Augmented Generation Scalable and Transparent
- LLMs as verification oracles for Solidity
- GSTM-HMU: Generative Spatio-Temporal Modeling for Human Mobility Understanding
- SMITE: Enhancing Fairness in LLMs through Optimal In-Context Example Selection via Dynamic Validation
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- No Labels Needed: Zero-Shot Image Classification with Collaborative Self-Learning
- Diversity Boosts AI-Generated Text Detection
- MAPO: Mixed Advantage Policy Optimization
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
- COLT: Enhancing Video Large Language Models with Continual Tool Usage
- TERAG: Token-Efficient Graph-Based Retrieval-Augmented Generation
- Database Normalization via Dual-LLM Self-Refinement
- Solving Math Word Problems Using Estimation Verification and Equation Generation
- Human-Annotated NER Dataset for the Kyrgyz Language
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- Multi-Hierarchical Feature Detection for Large Language Model Generated Text
- Advances in Large Language Models for Medicine
- Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
- MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service
- Anecdoctoring: Automated Red-Teaming Across Language and Place
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- The Persistence of Retracted Papers on Wikipedia
- Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation
- Accurate and Efficient Low-Rank Model Merging in Core Space
- ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
- CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
- MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responses
- EpiCache: Episodic KV Cache Management for Long Conversational Question Answering
- Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
- LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code
- Towards Open-Ended Discovery for Low-Resource NLP
- Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling
- Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
- AccessEval: Benchmarking Disability Bias in Large Language Models
- LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback
- Robustness of Neurosymbolic Reasoners on First-Order Logic Problems
- Adaptive Kernel Design for Bayesian Optimization Is a Piece of CAKE with LLMs
- How Persuasive is Your Context?
- Towards Provable Emergence of In-Context Reinforcement Learning
- ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification
- Improving Zero-shot Sentence Decontextualisation with Content Selection and Planning
- Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- Understanding Post-Training Structural Changes in Large Language Models
- Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs
- Specification-Aware Machine Translation and Evaluation for Purpose Alignment
- LIMI: Less is More for Agency
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- Qwen3-Omni Technical Report
- Probabilistic Token Alignment for Large Language Model Fusion
- CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages
- IDfRA: Self-Verification for Iterative Design in Robotic Assembly
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- Steering When Necessary: Flexible Steering Large Language Models with Backtracking
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- KAHAN: Knowledge-Augmented Hierarchical Analysis and Narration for Financial Data Narration
- SFT-TA: Supervised Fine-Tuned Agents in Multi-Agent LLMs for Automated Inductive Thematic Analysis
- Evolution of Concepts in Language Model Pre-Training
- AI Knows Best? The Paradox of Expertise, AI-Reliance, and Performance in Educational Tutoring Decision-Making Tasks
- Multi-level Diagnosis and Evaluation for Robust Tabular Feature Engineering with Large Language Models
- "Digital Camouflage": The LLVM Challenge in LLM-Based Malware Detection
- FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- Comparing RAG and GraphRAG for Page-Level Retrieval Question Answering on a Math Textbook
- RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
- Rethinking the Role of Text Complexity in Language Model Pretraining
- Federated Learning with Ad-hoc Adapter Insertions: The Case of Soft-Embeddings for Training Classifier-as-Retriever
- Towards Universal Debiasing for Language Models-based Tabular Data Generation
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
- Generative AI alone may not be enough: Evaluating AI Support for Learning Mathematical Proof
- LLM-Guided Co-Training for Text Classification
- Randomized Smoothing Meets Vision-Language Models
- Robust LLM Training Infrastructure at ByteDance
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
- Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
- Revisiting Vulnerability Patch Localization: An Empirical Study and LLM-Based Solution
- REFER: Mitigating Bias in Opinion Summarisation via Frequency Framed Prompting
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- It Depends: Resolving Referential Ambiguity in Minimal Contexts with Commonsense Knowledge
- Reward Hacking Mitigation using Verifiable Composite Rewards
- Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
- Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
- Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
- Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
- Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Sparse Multiview Open-Vocabulary 3D Detection
- DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
- Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
- Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- PolBiX: Detecting LLMs' Political Bias in Fact-Checking through X-phemisms
- LNE-Blocking: An Efficient Framework for Contamination Mitigation Evaluation on Large Language Models
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- Language Modeling with Learned Meta-Tokens
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models
- Semantic Representation Attack against Aligned Large Language Models
- CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
- Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
- Llama-Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens
- LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
- From Ground Trust to Truth: Disparities in Offensive Language Judgments on Contemporary Korean Political Discourse
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Reveal and Release: Iterative LLM Unlearning with Self-generated Data
- Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models
- Large Language Models in Operations Research: Methods, Applications, and Challenges
- Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agent Software
- A Scalable and Interoperable Platform for Transforming Building Information with Brick Ontology
- TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
- DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
- Training Larger Networks for Deep Reinforcement Learning
- Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
- Generative Large Language Models for Knowledge Representation: A Systematic Review of Concept Map Generation
- Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information Loss
- Synthetic bootstrapped pretraining
- Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs
- Charting trajectories of human thought using large language models
- A Taxonomy of Prompt Defects in LLM Systems
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Large Language Models as Universal Predictors? An Empirical Study on Small Tabular Datasets
- Jim137/qkan: v0.1.0
- UI-Level Evaluation of ALLaM 34B: Measuring an Arabic-Centric LLM via HUMAIN Chat
- Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
- Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
- State Space Models over Directed Graphs
- I, Robot? Exploring Ultra-Personalized AI-Powered AAC; an Autoethnographic Account
- ZERA: Zero-init Instruction Evolving Refinement Agent -- From Zero Instructions to Structured Prompts via Principle-based Optimization
- Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
- SAIL-VL2 Technical Report
- Learning the natural history of human disease with generative transformers
- An LLM-based multi-agent framework for agile effort estimation
- Synthetic Data Generation for Screen Time and App Usage
- Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
- Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings
- Do Large Language Models Understand Word Senses?
- CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
- CodeLSI: Leveraging Foundation Models for Automated Code Generation with Low-Rank Optimization and Domain-Specific Instruction Tuning
- ShortListing Model: A Streamlined SimplexDiffusion for Discrete Variable Generation
- Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
- Privacy Preserving In-Context-Learning Framework for Large Language Models
- Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
- Sequential Data Augmentation for Generative Recommendation
- Capturing Legal Reasoning Paths from Facts to Law in Court Judgments using Knowledge Graphs
- Risk Assessment and Security Analysis of Large Language Models
- Chinese Court Simulation with LLM-Based Agent System
- REAMS: Reasoning Enhanced Algorithm for Maths Solving
- Op-Fed: Opinion, Stance, and Monetary Policy Annotations on FOMC Transcripts Using Active Learning
- Crash Report Enhancement with Large Language Models: An Empirical Study
- Benchmarking ChatGPT and DeepSeek in April 2025: A Novel Dual Perspective Sentiment Analysis Using Lexicon-Based and Deep Learning Approaches
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
- Image Realness Assessment and Localization with Multimodal Features
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
- The Few-shot Dilemma: Over-prompting Large Language Models
- Automating Code Generation for Semiconductor Equipment Control from Developer Utterances with LLMs
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- Sparse Training Scheme for Multimodal LLM
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses
- Participatory AI: A Scandinavian Approach to Human-Centered AI
- FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- Analogy-Driven Financial Chain-of-Thought (AD-FCoT): A Prompting Approach for Financial Sentiment Analysis
- Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs
- Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
- Root Cause Analysis of Radiation Oncology Incidents Using Large Language Models
- Large Language Model-Based Automatic Formulation for Stochastic Optimization Models
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
- Exploring Training Data Attribution under Limited Access Constraints
- Contextualized Representation Learning for Effective Human-Object Interaction Detection
- InfoGain-RAG: Boosting Retrieval-Augmented Generation via Document Information Gain-based Reranking and Filtering
- Image-Seeking Intent Prediction for Cross-Device Product Search
- PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization
- Prompt Commons: Collective Prompting as Governance for Urban AI
- Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
- RAGs to Riches: RAG-like Few-shot Learning for Large Language Model Role-playing
- Pun Unintended: LLMs and the Illusion of Humor Understanding
- CBP-Tuning: Efficient Local Customization for Black-box Large Language Models
- Generative AI in Game Development: A Qualitative Research Synthesis
- Uncertainty in Authorship: Why Perfect AI Detection Is Mathematically Impossible
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- AssemMate: Graph-Based LLM for Robotic Assembly Assistance
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
- Graph-Enhanced Retrieval-Augmented Question Answering for E-Commerce Customer Support
- RAPTOR: A Foundation Policy for Quadrotor Control
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- Linguistic Neuron Overlap Patterns to Facilitate Cross-lingual Transfer on Low-resource Languages
- Anemoi: A Semi-Centralized Multi-agent System Based on Agent-to-Agent Communication MCP server from Coral Protocol
- Topic Coverage-based Demonstration Retrieval for In-Context Learning
- Quantifying Compositionality of Classic and State-of-the-Art Embeddings
- Policy Learning for Social Robot-Led Physiotherapy
- The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- Harnessing Optimization Dynamics for Curvature-Informed Model Merging
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
- Commenotes: Synthesizing Organic Comments to Support Community-Based Fact-Checking
- LoRALib: A Standardized Benchmark for Evaluating LoRA-MoE Methods
- The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- Differentially-private text generation degrades output language quality
- From Parameters to Performance: A Data-Driven Study on LLM Structure and Development
- Public Data Assisted Differentially Private In-Context Learning
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Verifying Computational Graphs in Production-Grade Distributed Machine Learning Frameworks
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- Contrastive Prompt Clustering for Weakly Supervised Semantic Segmentation
- Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
- RefactorCoderQA: Benchmarking LLMs for Multi-Domain Coding Question Solutions in Cloud and Edge Deployment
- WebSight: A Vision-First Architecture for Robust Web Agents
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
- A Discrepancy-Based Perspective on Dataset Condensation
- MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
- Compartmentalised Agentic Reasoning for Clinical NLI
- Benchmark of stylistic variation in LLM-generated texts
- Opening the Black Box: Interpretable LLMs via Semantic Resonance Architecture
- Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs
- Development of Automated Software Design Document Review Methods Using Large Language Models
- Beyond Token Limits: Assessing Language Model Performance on Long Text Classification
- Smart Trial: Evaluating the Use of Large Language Models for Recruiting Clinical Trial Participants via Social Media
- Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
- Unsupervised Hallucination Detection by Inspecting Reasoning Processes
- Enhancing LLM-based Specification Generation via Program Slicing and Logical Deletion
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- Fluent but Unfeeling: The Emotional Blind Spots of Language Models
- MetaLLMix : An XAI Aided LLM-Meta-learning Based Approach for Hyper-parameters Optimization
- DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones
- Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research
- Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
- CESRec: Constructing Pseudo Interactions for Sequential Recommendation via Conversational Feedback
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Medverse: A Universal Model for Full-Resolution 3D Medical Image Segmentation, Transformation and Enhancement
- Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
- Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment
- InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- How well can LLMs provide planning feedback in grounded environments?
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- LLMs as Agentic Cooperative Players in Multiplayer UNO
- IMDMR: An Intelligent Multi-Dimensional Memory Retrieval System for Enhanced Conversational AI
- Fast attention mechanisms: a tale of parallelism
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- Documents Are People and Words Are Items: A Psychometric Approach to Textual Data with Contextual Embeddings
- PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals
- Narrative-Guided Reinforcement Learning: A Platform for Studying Language Model Influence on Decision Making
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
- TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making
- A Role-Aware Multi-Agent Framework for Financial Education Question Answering with LLMs
- Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
- Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models
- Towards Scalable and Structured Spatiotemporal Forecasting
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution
- Efficient Decoding Methods for Language Models on Encrypted Data
- Rollout-LaSDI: Enhancing the long-term accuracy of Latent Space Dynamics
- Selective Induction Heads: How Transformers Select Causal Structures In Context
- Bias after Prompting: Persistent Discrimination in Large Language Models
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- Bringing Multi-Modal Multi-Task Federated Foundation Models to Education Domain: Prospects and Challenges
- Are Humans as Brittle as Large Language Models?
- Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
- Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Timing the Message: Language-Based Notifications for Time-Critical Assistive Settings
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrieval
- Causal Attention with Lookahead Keys
- Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
- M-BRe: Discovering Training Samples for Relation Extraction from Unlabeled Texts with Large Language Models
- Unleashing the True Potential of LLMs: A Feedback-Triggered Self-Correction with Long-Term Multipath Decoding
- In-Context Learning Enhanced Credibility Transformer
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- Reconstruction Alignment Improves Unified Multimodal Models
- HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
- SoK: Security and Privacy of AI Agents for Blockchain
- Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
- MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
- Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
- Intelligent Manufacturing Support: Specialized LLMs for Composite Material Processing and Equipment Operation
- Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
- Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
- From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
- O3Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
- Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
- Text-Trained LLMs Can Zero-Shot Extrapolate PDE Dynamics, Revealing a Three-Stage In-Context Learning Mechanism
- D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
- MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- IntrEx: A Dataset for Modeling Engagement in Educational Conversations
- Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial
- Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
- Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
- Natural Language-Programming Language Software Traceability Link Recovery Needs More than Textual Similarity
- LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
- Few-Shot Query Intent Detection via Relation-Aware Prompt Learning
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
- Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
- Prior Distribution and Model Confidence
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- Probabilistic operator learning: generative modeling and uncertainty quantification for foundation models of differential equations
- Masked Diffusion Language Models with Frequency-Informed Training
- Artificial intelligence for representing and characterizing quantum systems
- Scaling Law for Large-Scale Pre-Training Using Chaotic Time Series and Predictability in Financial Time Series
- L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
- Cloning a Conversational Voice AI Agent from Call Recording Datasets for Telesales
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Generative World Models of Tasks: LLM-Driven Hierarchical Scaffolding for Embodied Agents
- Memorization ≠ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?
- A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning
- Painting the market: generative diffusion models for financial limit order book simulation and forecasting
- Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
- Can VLMs Recall Factual Associations From Visual References?
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
- Finding your MUSE: Mining Unexpected Solutions Engine
- Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
- Shared Autonomy through LLMs and Reinforcement Learning for Applications to Ship Hull Inspections
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
- HumAIne-Chatbot: Real-Time Personalized Conversational AI via Reinforcement Learning
- An Empirical Study of Vulnerabilities in Python Packages and Their Detection
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Characterizing Fitness Landscape Structures in Prompt Engineering
- SMooGPT: Stylized Motion Generation using Large Language Models
- CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
- Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Leveraging LLM-Based Agents for Intelligent Supply Chain Planning
- Causality-guided Prompt Learning for Vision-language Models via Visual Granulation
- Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
- Boardwalk: Towards a Framework for Creating Board Games with LLMs
- Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
- The Physical Basis of Prediction: World Model Formation in Neural Organoids via an LLM-Generated Curriculum
- Hierarchical Federated Foundation Models over Wireless Networks for Multi-Modal Multi-Task Intelligence: Integration of Edge Learning with D2D/P2P-Enabled Fog Learning Architectures
- OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval
- CausalARC: Abstract Reasoning with Causal World Models
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- SOLD: SELFIES-based Objective-driven Latent Diffusion
- Language Models Do Not Follow Occam's Razor: A Benchmark for Inductive and Abductive Reasoning
- TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language Models
- Tabular foundation model for GEOAI benchmark problems BM/AirportSoilProperties/2/2025
- RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
- Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- FoMEMO: Towards Foundation Models for Expensive Multi-objective Optimization
- Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations
- GLARE: Agentic Reasoning for Legal Judgment Prediction
- Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
- FastMoE: A Fast Mixture-of-Expert Training System
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- OPRA-Vis: Visual Analytics System to Assist Organization-Public Relationship Assessment with Large Language Models
- The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
- Generative AI for Crystal Structures: A Review
- LLM Agents for Generating Microservice-based Applications: how complex is your specification?
- Perturbing the Derivative: Wild Refitting for Model-Free Evaluation of Machine Learning Models under Bregman Losses
- VASSO: Variance Suppression for Sharpness-Aware Minimization
- Benchmarking Large Language Models for Personalized Guidance in AI-Enhanced Learning
- ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
- A Survey of Embodied AI: From Simulators to Research Tasks
- Retrieval Enhanced Feedback via In-context Neural Error-book
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- Writing Polishment with Simile: Task, Dataset and A Neural Approach
- An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
- DivMerge: A divergence-based model merging method for multi-tasking
- Facts as Experts: Adaptable and Interpretable Neural Memory over\n Symbolic Knowledge
- Behavioral Fingerprinting of Large Language Models
- AutoDrive-R2: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
- Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
- HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers
- GridMind: LLMs-Powered Agents for Power System Analysis and Operations
- LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
- Deep Reinforcement Learning for Drone Route Optimization in Post-Disaster Road Assessment
- Automated Repair of C Programs Using Large Language Models
- Upcycling Candidate Tokens of Large Language Models for Query Expansion
- Extracting Training Data from Large Language Models
- Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain
- MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
- Relative Trajectory Balance is equivalent to Trust-PCL
- Inducing Faithfulness in Structured Reasoning via Counterfactual Sensitivity
- XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering
- KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
- Equivariant U-Shaped Neural Operators for the Cahn-Hilliard Phase-Field Model
- Iterative In-Context Learning to Enhance LLMs Abstract Reasoning: The Case-Study of Algebraic Tasks
- Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
- Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
- Question-to-Knowledge (Q2K): Multi-Agent Generation of Inspectable Facts for Product Mapping
- Generative Goal Modeling
- On the Alignment of Large Language Models with Global Human Opinion
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
- Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
- Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications
- Is All the Information in the Price? LLM Embeddings versus the EMH in Stock Clustering
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
- Whitening Sentence Representations for Better Semantics and Faster Retrieval
- Evaluating the Clinical Safety of LLMs in Response to High-Risk Mental Health Disclosures
- Can Large Language Models Master Complex Card Games?
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
- Breaking Barriers in Software Testing: The Power of AI-Driven Automation
- X-Troll: eXplainable Detection of State-Sponsored Information Operations Agents
- Supervised In-Context Fine-Tuning for Generative Sequence Labeling
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Feed Two Birds with One Scone: Exploiting Function-Space Regularization for Both OOD Robustness and ID Fine-Tuning Performance
- No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
- LLM Encoder vs. Decoder: Robust Detection of Chinese AI-Generated Text with LoRA
- OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
- Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
- Text Reinforcement for Multimodal Time Series Forecasting
- Towards Repository-Level Program Verification with Large Language Models
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- Neuro-Symbolic Predictive Process Monitoring
- Spotlighter: Revisiting Prompt Tuning from a Representative Mining View
- MedCOD: Enhancing English-to-Spanish Medical Translation of Large Language Models Using Enriched Chain-of-Dictionary Framework
- Inductive Biases and Variable Creation in Self-Attention Mechanisms
- Who Gets Left Behind? Auditing Disability Inclusivity in Large Language Models
- Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
- Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction
- PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement
- Can Multi-turn Self-refined Single Agent LMs with Retrieval Solve Hard Coding Problems?
- BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting
- COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
- SQL-of-Thought: Multi-agentic Text-to-SQL with Guided Error Correction
- KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
- LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
- Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
- Memory Limitations of Prompt Tuning in Transformers
- The Resurgence of GCG Adversarial Attacks on Large Language Models
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph
- Scalable Option Learning in High-Throughput Environments
- TimeCopilot
- A Modality-agnostic Multi-task Foundation Model for Human Brain Imaging
- Configuration Bugs Classification using LLMs and Encoders
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- CNN with large memory layers
- End-to-End Human Pose and Mesh Reconstruction with Transformers
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- Introducing LCOAI: A Standardized Economic Metric for Evaluating AI Deployment Costs
- CoComposer: LLM Multi-agent Collaborative Music Composition
- QZhou-Embedding Technical Report
- Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
- HealthProcessAI: A Technical Framework and Proof-of-Concept for LLM-Enhanced Healthcare Process Mining
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
- Artificial Intelligence in Drug Discovery: Applications and Techniques
- CXLAimPod: CXL Memory is all you need in AI era
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- Evaluation of Large Language Models for Anomaly Detection in Autonomous Vehicles
- From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
- Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study
- Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
- Summarize-Exemplify-Reflect: Data-driven Insight Distillation Empowers LLMs for Few-shot Tabular Classification
- Large Language Model Integration with Reinforcement Learning to Augment Decision-Making in Autonomous Cyber Operations
- Just-in-time and distributed task representations in language models
- Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization
- Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?
- Conditional Negative Sampling for Contrastive Learning of Visual Representations
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- GSTBench: A Benchmark Study on the Transferability of Graph Self-Supervised Learning
- An Agile Method for Implementing Retrieval Augmented Generation Tools in Industrial SMEs
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment
- MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
- The Application of Virtual Environments and Artificial Intelligence in Higher Education: Experimental Findings in Philosophy Teaching
- SemSR: Semantics aware robust Session-based Recommendations
- MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
- Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting
- Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction
- LLM Chatbot-Creation Approaches
- Tutorial on the Probabilistic Unification of Estimation Theory, Machine Learning, and Generative AI
- TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
- Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
- OneRec-V2 Technical Report
- Multi-Agent Penetration Testing AI for the Web
- Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- Governable AI: Provable Safety Under Extreme Threat Models
- BANG: Bridging Autoregressive and Non-autoregressive Generation with\n Large Scale Pretraining
- Turning Tabular Foundation Models into Graph Foundation Models
- Beyond Transcription: Mechanistic Interpretability in ASR
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- Visio-Verbal Teleimpedance Interface: Enabling Semi-Autonomous Control of Physical Interaction via Eye Tracking and Speech
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
- Uncertainty-Aware Collaborative System of Large and Small Models for Multimodal Sentiment Analysis
- Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
- Evaluating Language Model Reasoning about Confidential Information
- Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement
- Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey
- Ego-centric Predictive Model Conditioned on Hand Trajectories
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- PSO-Merging: Merging Models Based on Particle Swarm Optimization
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Interestingness First Classifiers
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Leveraging LLMs for Automated Translation of Legacy Code: A Case Study on PL/SQL to Java Transformation
- PersoNo: Personalised Notification Urgency Classifier in Mixed Reality
- A Scenario-Oriented Survey of Federated Recommender Systems: Techniques, Challenges, and Future Directions
- Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
- IELDG: Suppressing Domain-Specific Noise with Inverse Evolution Layers for Domain Generalized Semantic Segmentation
- Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
- Skill-based Explanations for Serendipitous Course Recommendation
- Robustness is Important: Limitations of LLMs for Data Fitting
- Orchid: Orchestrating Context Across Creative Workflows with Generative AI
- Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks
- ELIXIR: Efficient and LIghtweight model for eXplaIning Recommendations
- Surveying the Operational Cybersecurity and Supply Chain Threat Landscape when Developing and Deploying AI Systems
- CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- SIExVulTS: Sensitive Information Exposure Vulnerability Detection System using Transformer Models and Static Analysis
- Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- Autoregressive Universal Video Segmentation Model
- Evaluating the Evaluators: Are readability metrics good measures of readability?
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Measuring and Improving Consistency in Pretrained Language Models
- SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
- ZeST: an LLM-based Zero-Shot Traversability Navigation for Unknown Environments
- TEASEL: A Transformer-Based Speech-Prefixed Language Model
- M-LLM3REC: A Motivation-Aware User-Item Interaction Framework for Enhancing Recommendation Accuracy with LLMs
- Retrieval-Augmented Review Generation for Poisoning Recommender Systems
- An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
- Deep Learning-Enabled Supercritical Flame Simulation at Detailed Chemistry and Real-Fluid Accuracy Towards Trillion-Cell Scale
- Interleaving Large Language Models for Compiler Testing
- Novel Approaches to Artificial Intelligence Development Based on the Nearest Neighbor Method
- Neural Symbolic Regression that Scales
- ReflectivePrompt: Reflective evolution in autoprompting algorithms
- Directed Beam Search: Plug-and-Play Lexically Constrained Language Generation
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- CASP: An evaluation dataset for formal verification of C code
- LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordination
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- A Study of Privacy-preserving Language Modeling Approaches
- Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Bias Mitigation Agent: Optimizing Source Selection for Fair and Balanced Knowledge Retrieval
- Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
- UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
- Automatic Prompt Optimization with Prompt Distillation
- Language and Experience: A Computational Model of Social Learning in Complex Tasks
- Revisiting associative recall in modern recurrent models
- Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
- DIO: Refining Mutual Information and Causal Chain to Enhance Machine Abstract Reasoning Ability
- COMET-poly: Machine Translation Metric Grounded in Other Candidates
- Skeptik: A Hybrid Framework for Combating Potential Misinformation in Journalism
- Cognitive Agents Powered by Large Language Models for Agile Software Project Management
- Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- From BERT to LLMs: Comparing and Understanding Chinese Classifier Prediction in Language Models
- Type-Compliant Adaptation Cascades: Adapting Programmatic LM Workflows to Data
- What they do when in doubt: a study of inductive biases in seq2seq learners
- Unraveling the cognitive patterns of Large Language Models through module communities
- Leveraging Large Language Models for Accurate Sign Language Translation in Low-Resource Scenarios
- Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- Named Entity Recognition of Historical Text via Large Language Model
- Multimodal Conditionality for Natural Language Generation
- Contrastive Representation Learning for Exemplar-Guided Paraphrase Generation
- VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
- Self-Guided Function Calling in Large Language Models via Stepwise Experience Recall
- SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
- OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs
- Automatic Graph Partitioning for Very Large-scale Deep Learning
- QMUL-SDS at SCIVER: Step-by-Step Binary Classification for Scientific Claim Verification
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
- LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
- Select to Know: An Internal-External Knowledge Self-Selection Framework for Domain-Specific Question Answering
- Adversarial Attacks against Neural Ranking Models via In-Context Learning
- Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
- Identifying and Answering Questions with False Assumptions: An Interpretable Approach
- Tensorized Multi-Task Learning for Personalized Modeling of Heterogeneous Individuals with High-Dimensional Data
- Frequency-adaptive tensor neural networks for high-dimensional multi-scale problems
- Mapping the Course for Prompt-based Structured Prediction
- LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- A Theory of Information, Variation, and Artificial Intelligence
- Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
- Investigation of the Inter-Rater Reliability between Large Language Models and Human Raters in Qualitative Analysis
- GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
- Collective Intelligence for Deep Learning: A Survey of Recent Developments
- Scaled Signed Averaging Improves In-Context and Early Learning Benchmark Performance in Small Transformers
- ELATE: Evolutionary Language model for Automated Time-series Engineering
- Understanding Data Influence with Differential Approximation
- Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
- PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments
- Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
- In2x at WMT25 Translation Task
- Automated Optimization Modeling through Expert-Guided Large Language Model Reasoning
- Can AI Have a Personality? Prompt Engineering for AI Personality Simulation: A Chatbot Case Study in Gender-Affirming Voice Therapy Training
- Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
- Adaptive Interpolating Quantum Transform: A Quantum-Native Framework for Efficient Transform Learning
- The Prompting Brain: Neurocognitive Markers of Expertise in Guiding Large Language Models
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- In-Context Iterative Policy Improvement for Dynamic Manipulation
- Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
- Data Distillation for Text Classification
- Pixels to Play: A Foundation Model for 3D Gameplay
- Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
- Comparing energy consumption and accuracy in text classification inference
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Prompt Orchestration Markup Language
- Revisiting RAG Ensemble: A Theoretical and Mechanistic Analysis of Multi-RAG System Collaboration
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- Can Large Language Models (LLMs) Describe Pictures Like Children? A Comparative Corpus Study
- Generics and Default Reasoning in Large Language Models
- subCellSAM: Zero-Shot (Sub-)Cellular Segmentation for Hit Validation in Drug Discovery
- Knowledge Graph Completion for Action Prediction on Situational Graphs -- A Case Study on Household Tasks
- In-Context Decision Making for Optimizing Complex AutoML Pipelines
- Interpreting the Interpreter: Can We Model post-ECB Conferences Volatility with LLM Agents?
- Towards a Larger Model via One-Shot Federated Learning on Heterogeneous Client Models
- Equinox: Holistic Fair Scheduling in Serving Large Language Models
- Scalable Scientific Interest Profiling Using Large Language Models
- Genuine multipartite entanglement verification with convolutional neural networks
- Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
- ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- Graph Concept Bottleneck Models
- Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
- X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
- DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
- AI Agents for Photonic Integrated Circuit Design Automation
- Exploring Story Generation with Multi-task Objectives in Variational Autoencoders
- Reinforced Context Order Recovery for Adaptive Reasoning and Planning
- CardAIc-Agents: A Multimodal Framework with Hierarchical Adaptation for Cardiac Care Support
- REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks
- When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
- Wavy Transformer
- MemorySim: An RTL-level, timing accurate simulator model for the Chisel ecosystem
- Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
- Breaking Language Barriers: Equitable Performance in Multilingual Language Models
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
- Learning In-context \pmbn-grams with Transformers: Sub-\pmbn-grams Are Near-stationary Points
- Informative data visualization with raincloud plots in JASP
- Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction
- Next Visual Granularity Generation
- TASER: Table Agents for Schema-guided Extraction and Recommendation
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
- Feature Request Analysis and Processing: Tasks, Techniques, and Trends
- Non-Interactive Symbolic-Aided Chain-of-Thought for Logical Reasoning
- GALA: Can Graph-Augmented Large Language Model Agentic Workflows Elevate Root Cause Analysis?
- EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
- Mantis: A Simulation-Grounded Foundation Model for Disease Forecasting
- ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression
- Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
- Cost-Aware Contrastive Routing for LLMs
- Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark
- Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Research on Conversational Recommender System Considering Consumer Types
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- Limitation Learning: Catching Adverse Dialog with GAIL
- Assessing User Privacy Leakage in Synthetic Packet Traces: An Attack-Grounded Approach
- Controlling Multimodal LLMs via Reward-guided Decoding
- TinyTim: A Family of Language Models for Divergent Generation
- CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
- When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
- Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
- UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
- Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
- Understanding in Artificial Intelligence
- ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
- Human-in-the-Loop Systems for Adaptive Learning Using Generative AI
- WIP: Leveraging LLMs for Enforcing Design Principles in Student Code: Analysis of Prompting Strategies and RAG
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- Improving Text Style Transfer using Masked Diffusion Language Models with Inference-time Scaling
- Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks
- Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
- STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- Searching for Privacy Risks in LLM Agents via Simulation
- SSRL: Self-Search Reinforcement Learning
- Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
- IBEX: Information-Bottleneck-EXplored Coarse-to-Fine Molecular Generation under Limited Data
- FROGENT: An End-to-End Full-process Drug Design Agent
- Hypercomplex Prompt-aware Multimodal Recommendation
- Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Detecting Hate Speech with GPT-3
- Advancing Autonomous Incident Response: Leveraging LLMs and Cyber Threat Intelligence
- SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
- STEP: Stepwise Curriculum Learning for Context-Knowledge Fusion in Conversational Recommendation
- MSRS: Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- Inductive Bias Extraction and Matching for LLM Prompts
- Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
- Dirichlet Pruning for Neural Network Compression
- Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
- Chain-of-Query: Unleashing the Power of LLMs in SQL-Aided Table Understanding via Multi-Agent Collaboration
- LingVarBench: Benchmarking LLM for Automated Named Entity Recognition in Structured Synthetic Spoken Transcriptions
- Can Transformers Break Encryption Schemes via In-Context Learning?
- Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Pre-trained Transformer-models using chronic invasive electrophysiology for symptom decoding without patient-individual training
- Constrained Decoding of Diffusion LLMs with Context-Free Grammars
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- Data-Driven Discovery of Interpretable Kalman Filter Variants through Large Language Models and Genetic Programming
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- Teaching LLMs to Speak Spectroscopy
- Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
- Adoption of Explainable Natural Language Processing: Perspectives from Industry and Academia on Practices and Challenges
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- The PacifAIst Benchmark:Would an Artificial Intelligence Choose to Sacrifice Itself for Human Safety?
- ReqInOne: A Large Language Model-Based Agent for Software Requirements Specification Generation
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round Annotation
- EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- CS-Agent: LLM-based Community Search via Dual-agent Collaboration
- Decentralized Rank Scheduling for Energy-Constrained Multi-Task Federated Fine-Tuning in Edge-Assisted IoV Networks
- Learning Facts at Scale with Active Reading
- Security Analysis of ChatGPT: Threats and Privacy Risks
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
- Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
- Wide Neural Networks Forget Less Catastrophically
- The Othello AI Arena: Evaluating Intelligent Systems Through Limited-Time Adaptation to Unseen Boards
- Dual-stream Network for Visual Recognition
- Scaling Up Active Testing to Large Language Models
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- Reveal-Bangla: A Dataset for Cross-Lingual Multi-Step Reasoning Evaluation
- Compass-Thinker-7B Technical Report
- InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
- A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
- Optimal ANN-SNN Conversion for Fast and Accurate Inference in Deep Spiking Neural Networks
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
- Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
- Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
- TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
- GreenTEA: Gradient Descent with Topic-modeling and Evolutionary Auto-prompting
- First-Generation Inference Accelerator Deployment at Facebook
- Prompt-and-Check: Using Large Language Models to Evaluate Communication Protocol Compliance in Simulation-Based Training
- Special-Character Adversarial Attacks on Open-Source Language Model
- DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- Superclass-Guided Representation Disentanglement for Spurious Correlation Mitigation
- OmniLLP: Enhancing LLM-based Log Level Prediction with Context-Aware Retrieval
- LyS at SemEval 2025 Task 8: Zero-Shot Code Generation for Tabular QA
- Leveraging Large Language Models for Rare Disease Named Entity Recognition
- MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
- Shaping the future of education: a cluster analysis of generative AI’s transformative impact
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- 3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
- SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling
- Data-Efficient Biomedical In-Context Learning: A Diversity-Enhanced Submodular Perspective
- Holdout-Based Fidelity and Privacy Assessment of Mixed-Type Synthetic Data
- Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- Grouped Speculative Decoding for Autoregressive Image Generation
- LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval
- AIS-LLM: A Unified Framework for Maritime Trajectory Prediction, Anomaly Detection, and Collision Risk Assessment with Explainable Forecasting
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- Autonomous Navigation of Cloud-Controlled Quadcopters in Confined Spaces Using Multi-Modal Perception and LLM-Driven High Semantic Reasoning
- MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
- CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
- Large Language Models for Subjective Language Understanding: A Survey
- Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation
- Can You Trick the Grader? Adversarial Persuasion of LLM Judges
- A Spin Glass Characterization of Neural Networks
- Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
- Prompt Tuning for Few-Shot Continual Learning Named Entity Recognition
- Propagation Tree Is Not Deep: Adaptive Graph Contrastive Learning Approach for Rumor Detection
- Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Semantic Alignment
- Enhancing Rumor Detection Methods with Propagation Structure Infused Language Model
- LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
- Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
- Towards Real-World Rumor Detection: Anomaly Detection Framework with Graph Supervised Contrastive Learning
- Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
- Generative AI for Strategic Plan Development
- The ReQAP System for Question Answering over Personal Information
- ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Many-Turn Jailbreaking
- Integrating Rules and Semantics for LLM-Based C-to-Rust Translation
- Bridging Classical and Quantum Computing for Next-Generation Language Models
- Confidence Estimation for Text-to-SQL in Large Language Models
- Diffeomorphic Neural Operator Learning
- gpt-oss-120b & gpt-oss-20b Model Card
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- Position: Ideas Should be the Center of Machine Learning Research
- A game-changer for qualitative research: artificial intelligence as an efficient tool for analyzing student conceptions about microplastics
- ChatGPT Relies More Heavily on Consonants Than on Vowels to Recognize Words
- Exploring a biocentric LLM-based assistant in environmental decision-making with more-than-human representation of the Tagus Estuary
- ToPolyAgent: AI agents for coarse-grained bead-spring topological polymer simulations
- Cyberbullying Detection via Aggression-Enhanced Prompting
- Aligning Effective Tokens with Video Anomaly in Large Language Models
- Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
- SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
- LLM Serving Optimization with Variable Prefill and Decode Lengths
- PanelTR: Zero-Shot Table Reasoning Framework Through Multi-Agent Scientific Discussion
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- When a Paper Has 1000 Authors: Rethinking Citation Metrics in the Era of LLMs
- LinguaFluid: Language Guided Fluid Control via Semantic Rewards in Reinforcement Learning
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- Role of Large Language Models and Retrieval-Augmented Generation for Accelerating Crystalline Material Discovery: A Systematic Review
- Scaling Personality Control in LLMs with Big Five Scaler Prompts
- LLMs for Resource Allocation: A Participatory Budgeting Approach to Inferring Preferences
- RTTC: Reward-Guided Collaborative Test-Time Compute
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
- A Framework for Inherently Safer AGI through Language-Mediated Active Inference
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety
- Streamlining Admission with LOR Insights: AI-Based Leadership Assessment in Online Master's Program
- GRAIL:Learning to Interact with Large Knowledge Graphs for Retrieval Augmented Reasoning
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
- Simulating LLM training workloads for heterogeneous compute and network infrastructure
- Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
- CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
- AI-assisted JSON Schema Creation and Mapping
- Posterior-GRPO: Rewarding Reasoning Processes in Code Generation
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- Attention Basin: Why Contextual Position Matters in Large Language Models
- Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
- Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes
- Generative AI for Object-Oriented Programming: Writing the Right Code and Reasoning the Right Logic
- Situated Epistemic Infrastructures: A Diagnostic Framework for Post-Coherence Knowledge
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- An Effective Approach for Node Classification in Textual Graphs
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- Root Cause Analysis Training for Healthcare Professionals With AI-Powered Virtual Simulation: A Proof-of-Concept
- Open Scene Graphs for Open-World Object-Goal Navigation
- X-SAM: From Segment Anything to Any Segmentation
- RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
- P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
- A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
- Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
- A Unified Pre-training Framework for Conversational AI
- Low-Precision Hardware Architectures Meet Recommendation Model Inference at Scale
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- TRAIL: Joint Inference and Refinement of Knowledge Graphs with Large Language Models
- GFocal: A Global-Focal Neural Operator for Solving PDEs on Arbitrary Geometries
- StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
- PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
- Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
- LUST: A Multi-Modal Framework with Hierarchical LLM-based Scoring for Learned Thematic Significance Tracking in Multimedia Content
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- Mockingbird: How does LLM perform in general machine learning tasks?
- Large Language Model's Multi-Capability Alignment in Biomedical Domain
- KG-Augmented Executable CoT for Mathematical Coding
- Fine-tuning for Better Few Shot Prompting: An Empirical Comparison for Short Answer Grading
- Graph Representation Learning with Massive Unlabeled Data for Rumor Detection
- T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion
- DP-GPT4MTS: Dual-Prompt Large Language Model for Textual-Numerical Time Series Forecasting
- Hierarchical Text Classification Using Black Box Large Language Models
- Guided Navigation in Knowledge-Dense Environments: Structured Semantic Exploration with Guidance Graphs
- AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
- Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts
- ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation
- AquaChat++: LLM-Assisted Multi-ROV Inspection for Aquaculture Net Pens with Integrated Battery Management and Thruster Fault Tolerance
- SVC 2025: the First Multimodal Deception Detection Challenge
- Neuro-MoBRE: Exploring Multi-subject Multi-task Intracranial Decoding via Explicit Heterogeneity Resolving
- Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks
- Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
- A Foundational Multi-Modal Model for Few-Shot Learning
- Enhancing Serendipity Recommendation System by Constructing Dynamic User Knowledge Graphs with Large Language Models
- StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback
- Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
- An Entity Linking Agent for Question Answering
- MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- FairLangProc: A Python package for fairness in NLP
- DiWA: Diffusion Policy Adaptation with World Models
- SoilNet: A Multimodal Multitask Model for Hierarchical Classification of Soil Horizons
- SAGE-HLS: Syntax-Aware AST-Guided LLM for High-Level Synthesis Code Generation
- EditGarment: An Instruction-Based Garment Editing Dataset Constructed with Automated MLLM Synthesis and Semantic-Aware Evaluation
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction using Enhanced Triple Extraction
- Variety Is the Spice of Life: Detecting Misinformation with Dynamic Environmental Representations
- When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
- From Legacy to Standard: LLM-Assisted Transformation of Cybersecurity Playbooks into CACAO Format
- Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
- LECTOR: LLM-Enhanced Concept-based Test-Oriented Repetition for Adaptive Spaced Learning
- Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
- ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
- A System Model Generation Benchmark from Natural Language Requirements
- Current State in Privacy-Preserving Text Preprocessing for Domain-Agnostic NLP
- From Text to Trajectories: GPT-2 as an ODE Solver via In-Context
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- Multimodal Human-Intent Modeling for Contextual Robot-to-Human Handovers of Arbitrary Objects
- CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
- Knowledge Transfer via Pre-training for Recommendation: A Review and Prospect
- Thinking with Nothinking Calibration: A New In-Context Learning Paradigm in Reasoning Large Language Models
- CTTS: Collective Test-Time Scaling
- Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment
- GrandJury: A Collaborative Machine Learning Model Evaluation Protocol for Dynamic Quality Rubrics
- Tricks and Plug-ins for Gradient Boosting with Transformers
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- AutoGeTS: Knowledge-based Automated Generation of Text Synthetics for Improving Text Classification
- Neural Scaling Laws Surpass Chemical Accuracy for the Many-Electron Schrödinger Equation
- OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling
- CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
- Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- PCREQ: Automated Inference of Compatible Requirements for Python Third-party Library Upgrades
- Evaluating Position Bias in Large Language Model Recommendations
- FinCPRG: A Bidirectional Generation Pipeline for Hierarchical Queries and Rich Relevance in Financial Chinese Passage Retrieval
- TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
- Prompting Large Language Models to Detect Dementia Family Caregivers
- Whispering Agents: An Event-driven Covert Communication Protocol For the Internet of Agents
- A Survey on Data Security in Large Language Models
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- A Methodological Framework for LLM-Based Mining of Software Repositories
- PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs
- ProCut: LLM Prompt Compression via Attribution Estimation
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
- MLP Memory: A Retriever-Pretrained Memory for Large Language Models
- Social Media Information Operations
- Getting out of the Big-Muddy: Escalation of Commitment in LLMs
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- The potential and limitations of large language models for automatic classification of teachers' motivational messages in educational research
- From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models
- LLM-Assisted Model-Based Fuzzing of Protocol Implementations
- RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
- How Does Controllability Emerge In Language Models During Pretraining?
- A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agents
- Fast and scalable retrosynthetic planning with a transformer neural network and speculative beam search
- From Query to Logic: Ontology-Driven Multi-Hop Reasoning in LLMs
- MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs
- ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
- Kronos: A Foundation Model for the Language of Financial Markets
- How Far Are LLMs from Symbolic Planners? An NLP-Based Perspective
- Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities
- Win-k: Improved Membership Inference Attacks on Small Language Models
- Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
- Multi-Operator Few-Shot Learning for Generalization Across PDE Families
- T2S: Tokenized Skill Scaling for Lifelong Imitation Learning
- Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
- CSIRO-LT at SemEval-2025 Task 11: Adapting LLMs for Emotion Recognition for Multiple Languages
- BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation
- Disaggregated Health Data in LLMs: Evaluating Data Equity in the Context of Asian American Representation
- A Note on Code Quality Score: LLMs for Maintainable Large Codebases
- Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education
- Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
- SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
- How LLMs are Shaping the Future of Virtual Reality
- PaPaformer: Language Model from Pre-trained Parallel Paths
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Automated Type Annotation in Python Using Large Language Models
- Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
- FeatureCuts: Feature Selection for Large Data by Optimizing the Cutoff
- Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
- DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
- ITUNLP at SemEval-2025 Task 8: Question-Answering over Tabular Data: A Zero-Shot Approach using LLM-Driven Code Generation
- MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- Lessons from complex systems science for AI governance
- ITDR: An Instruction Tuning Dataset for Enhancing Large Language Models in Recommendations
- GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
- TriP-LLM: A Tri-Branch Patch-wise Large Language Model Framework for Time-Series Anomaly Detection
- CFDagent: A Language-Guided, Zero-Shot Multi-Agent System for Complex Flow Simulation
- Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
- Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
- TT-Rec: Tensor Train Compression for Deep Learning Recommendation Models
- On Extending NLP Techniques from the Categorical to the Latent Space: KL Divergence, Zipf's Law, and Similarity Search
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- How Far Are AI Scientists from Changing the World?
- DynaSwarm: Dynamically Graph Structure Selection for LLM-based Multi-agent System
- Personalized Education with Ranking Alignment Recommendation
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- LLM4Rail: An LLM-Augmented Railway Service Consulting Platform
- Causal Reasoning in Pieces: Modular In-Context Learning for Causal Discovery
- Enabling Few-Shot Alzheimer's Disease Diagnosis on Biomarker Data with Tabular LLMs
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
- ChatVis: Large Language Model Agent for Generating Scientific Visualizations
- KLLM: Fast LLM Inference with K-Means Quantization
- Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
- TR-PTS: Task-Relevant Parameter and Token Selection for Efficient Tuning
- Beyond Natural Language Plans: Structure-Aware Planning for Query-Focused Table Summarization
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- OFCnetLLM: Large Language Model for Network Monitoring and Alertness
- GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
- Multilingual Political Views of Large Language Models: Identification and Steering
- Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- On the Definition of Intelligence
- PATENTWRITER: A Benchmarking Study for Patent Drafting with LLMs
- DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
- An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem
- User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
- A Foundation Model for Material Fracture Prediction
- Exploring In-Context Learning for Frame-Semantic Parsing
- Context-aware Rotary Position Embedding
- Next Tokens Denoising for Speech Synthesis
- Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items
- Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
- Valuing Time in Silicon: Can Large Language Models Replicate Human Value of Travel Time
- TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction
- Generative Recommendation with Semantic IDs: A Practitioner's Handbook
- When Truthful Representations Flip Under Deceptive Instructions?
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
- ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
- Predicting Microbial Ontology and Pathogen Risk from Environmental Metadata with Large Language Models
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Low-Regret Active learning
- Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences
- Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
- MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning
- Domain Generalization and Adaptation in Intensive Care with Anchor Regression
- AU-LLM: Micro-Expression Action Unit Detection via Enhanced LLM-Based Feature Fusion
- Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
- MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
- RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- AI Literacy as a Key Driver of User Experience in AI-Powered Assessment: Insights from Socratic Mind
- Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
- Ethical Classification of Non-Coding Contributions in Open-Source Projects via Large Language Models
- Transmission With Machine Language Tokens: A Paradigm for Task-Oriented Agent Communication
- MindChat: Enhancing BCI Spelling with Large Language Models in Realistic Scenarios
- Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour
- Shapley Uncertainty in Natural Language Generation
- AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini
- Knowledge Editing for Multi-Hop Question Answering Using Semantic Analysis
- Neural Autoregressive Modeling of Brain Aging
- CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
- Memorization in Fine-Tuned Large Language Models
- Language Models as Few-Shot Learner for Task-Oriented Dialogue Systems
- TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
- Prescriptive Agents based on RAG for Automated Maintenance (PARAM)
- Leveraging Open-Source Large Language Models for Clinical Information Extraction in Resource-Constrained Settings
- Latent Inter-User Difference Modeling for LLM Personalization
- METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
- First Hallucination Tokens Are Different from Conditional Ones
- Regularizing Subspace Redundancy of Low-Rank Adaptation
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- TokenPose: Learning Keypoint Tokens for Human Pose Estimation
- Ontology-Enhanced Knowledge Graph Completion using Large Language Models
- Cut the CARP: Fishing for zero-shot story evaluation
- A fast memoryless predictive algorithm in a chain of recurrent neural networks
- Enhancing Hallucination Detection via Future Context
- Beyond Class Tokens: LLM-guided Dominant Property Mining for Few-shot Classification
- Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
- Provable In-Context Learning of Nonlinear Regression with Transformers
- GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- Foundation model [wikipedia]
- GPT-2 [wikipedia]
- GPT-3 [wikipedia]
- GPT-4 [wikipedia]
- Generative AI [wikipedia]
- Generative pre-trained transformer [wikipedia]
- History of natural language processing [wikipedia]
- Jared Kaplan [wikipedia]
- Language model [wikipedia]
- Large language model [wikipedia]
- List of large language models [wikipedia]
- Narrative intelligence [wikipedia]
- Neural scaling law [wikipedia]
- Neuro-symbolic AI [wikipedia]
- Olin College [wikipedia]
- Products and applications of OpenAI [wikipedia]
- Prompt engineering [wikipedia]
- Synthetic media [wikipedia]
- Winograd schema challenge [wikipedia]
- Wu Dao [wikipedia]
- Timeline of artificial intelligence [wikipedia]
- AI alignment [wikipedia]
- Amanda Askell [wikipedia]
- Byte-pair encoding [wikipedia]
- DALL-E [wikipedia]
Discussions
- GPT-3: Language Models Are Few-Shot Learners [hn, 431 points, 201 comments]
- this has to be the paper that, looking back at the authors from today's context, you're like, oh damn all those people really did work together arxiv.org/abs/2005.14165 [bsky, 31 points, 1 comments]
- also umm they do learn unless you define learning as requiring to be implemented thru a change in the weights arxiv.org/abs/2005.14165 [bsky, 5 points, 2 comments]
- no, it really did happen: arxiv.org/abs/2005.14165 *nobody* expected LLMs to be able to answer arbitrary queries from just being scaled up [bsky, 4 points, 1 comments]
- Yes I think we (and many others, e.g. fig 1 of arxiv.org/abs/2005.14165) interpret it in broadly that way! There's a sorta natural division of inner loop = in-context adaptation, outer loop = training [bsky, 2 points, 1 comments]
- They originally just promised few-shot, I think. But basically yeah. arxiv.org/abs/2005.14165 [bsky, 2 points, 0 comments]
- OpenAI claims the NYT is barred by the statute of limitations bc it should have known of the alleged directly infringing activity by the time this academic article that does not mention the NY Times w [bsky, 2 points, 1 comments]
- Language Models Are Few-Shot Learners [hn, 2 points, 0 comments]
- Language models are few-shot learners (2020) [hn, 1 points, 0 comments]
- GPT [bsky, 0 points, 0 comments]
- GPT3: Unsupervised pre-training + a few* examples is all you need. *Up to 5 (Conversational QA) - 50 examples (Winogrande, PhysicalQA, TriviaQA) https://arxiv.org/abs/2005.14165 [bsky, 0 points, 1 comments]
- many such cases arxiv.org/abs/2005.14165 [bsky, 0 points, 1 comments]
- 🎧 EP016: GPT-3 Learns From Examples Without Retraining 📄 GPT-3 🔗 https://arxiv.org/abs/2005.14165 🟢 https://podcasters.spotify.com/pod/show/yun-wu/episodes/EP016-GPT-3-Learns-From-Examples-Without [bsky, 0 points, 0 comments]
- 1️⃣ GPT-3 (mayo 2020): los LLMs dejan de ser una curiosidad académica para convertirse en una tecnología disruptiva. 🔗 Paper: arxiv.org/abs/2005.14165 Demostró capacidades de "few-shot learning". Gen [bsky, 0 points, 1 comments]
- Spannend ist die Diskussion des im Paper "Language Models are Few-Shot Learners" (arxiv.org/abs/2005.14165) beschriebenen In-Kontext- oder Meta-Lernens von Transformern wie #ChatGPT, bei dem im Prompt [bsky, 0 points, 1 comments]
- RT @hardmaru: GPT-3: Language Models are Few-Shot Learners, by @notTomBrown et al. “We train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse [bsky, 0 points, 0 comments]
Related