Character-level Convolutional Networks for Text Classification
2015/09/04 by Xiang Zhang, Junbo Zhao, Zhang, Xiang +3 · 1 voice · 354 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text and Document Classification Technologies #Topic Modeling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1509.01626
An early version of this work entitled "Text Understanding from Scratch" was posted in Feb 2015 as arXiv:1502.01710. The present paper has considerably more experimental results and a rewritten introduction, Advances in Neural Information Processing Systems 28 (NIPS 2015)
openalex publication_date 2015/09/04 · arxiv published 2015/09/04 · arxiv created 2016/04/04 · arxiv updated 2016/04/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
This article offers an empirical exploration on the use of character-level convolutional networks (ConvNets) for text classification. We constructed several large-scale datasets to show that character-level convolutional networks could achieve state-of-the-art or competitive results. Comparisons are offered against traditional models such as bag of words, n-grams and their TFIDF variants, and deep learning models such as word-based ConvNets and recurrent neural networks.
Citations
Cited by
- Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
- Automated Numerical Stability Analysis of Deep Learning Operators
- When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness
- Clust-PSI-PFL: A Population Stability Index Approach for Clustered Non-IID Personalized Federated Learning
- HATS: High-Accuracy Triple-Set Watermarking for Large Language Models
- MAGIC: Achieving Superior Model Merging via Magnitude Calibration
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
- Calibrating Transformer Attention via Task-Space Sensitivity Feedback
- From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
- Convolutional Lie Operator for Sentence Classification
- PerProb: Indirectly Evaluating Memorization in Large Language Models
- Measuring Uncertainty Calibration
- Reducing Label Dependency in Human Activity Recognition with Wearables: From Supervised Learning to Novel Weakly Self-Supervised Approaches
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- THeGAU: Type-Aware Heterogeneous Graph Autoencoder and Augmentation
- Interpreto: An Explainability Library for Transformers
- Text2Graph: Combining Lightweight LLMs and GNNs for Efficient Text Classification in Label-Scarce Scenarios
- Improving Multi-Class Calibration through Normalization-Aware Isotonic Techniques
- Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
- TopiCLEAR: Topic extraction by CLustering Embeddings with Adaptive dimensional Reduction
- LOCUS: A System and Method for Low-Cost Customization for Universal Specialization
- PrivCode: When Code Generation Meets Differential Privacy
- What Language is This? Ask Your Tokenizer
- Technical Report on Text Dataset Distillation
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- One Swallow Does Not Make a Summer: Understanding Semantic Structures in Embedding Spaces
- Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
- SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
- Masks Can Be Distracting: On Context Comprehension in Diffusion Language Models
- Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
- Geometry of Decision Making in Language Models
- When +1% Is Not Enough: A Paired Bootstrap Protocol for Evaluating Small Improvements
- MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- Re-Key-Free, Risky-Free: Adaptable Model Usage Control
- A Lightweight Approach to Detection of AI-Generated Texts Using Stylometric Features
- PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
- DelTriC: A Novel Clustering Method with Accurate Outlier
- Attention-Guided Feature Fusion (AGFF) Model for Integrating Statistical and Semantic Features in News Text Classification
- Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
- ILoRA: Federated Learning with Low-Rank Adaptation for Heterogeneous Client Aggregation
- GPS: General Per-Sample Prompter
- Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
- Mitigating Label Length Bias in Large Language Models
- SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
- SnapAudit: Active Auditing of Differentially Private In-Context Learning via Snapshot-Based Simulation
- RegionMarker: A Region-Triggered Semantic Watermarking Framework for Embedding-as-a-Service Copyright Protection
- Generalization Bounds for Semi-supervised Matrix Completion with Distributional Side Information
- Private Zeroth-Order Optimization with Public Data
- Advanced Black-Box Tuning of Large Language Models with Limited API Calls
- AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
- Zero-Order Sharpness-Aware Minimization
- iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
- Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks
- Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
- Graph Representation-based Model Poisoning on the Heterogeneous Internet of Agents
- CAE: Character-Level Autoencoder for Non-Semantic Relational Data Grouping
- Federated Stochastic Minimax Optimization under Heavy-Tailed Noises
- Differentially Private In-Context Learning with Nearest Neighbor Search
- On Joint Regularization and Calibration in Deep Ensembles
- Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
- Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction
- Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
- Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
- Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
- Scalable Utility-Aware Multiclass Calibration
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
- A high-capacity linguistic steganography based on entropy-driven rank-token mapping
- Simple Denoising Diffusion Language Models
- Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
- SALSA: Single-pass Autoregressive LLM Structured Classification
- Power to the Clients: Federated Learning in a Dictatorship Setting
- Frequentist Validity of Epistemic Uncertainty Estimators
- DictPFL: Efficient and Private Federated Learning on Encrypted Gradients
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
- Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
- Latent-Augmented Discrete Diffusion Models
- Fighter: Unveiling the Graph Convolutional Nature of Transformers in Time Series Modeling
- Rotation, Scale, and Translation Resilient Black-box Fingerprinting for Intellectual Property Protection of EaaS Models
- The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
- CurLL: A Developmental Framework to Evaluate Continual Learning in Language Models
- Market-Driven Subset Selection for Budgeted Training
- FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management
- Investigating Large Language Models' Linguistic Abilities for Text Preprocessing
- Backdoor Collapse: Eliminating Unknown Threats via Known Backdoor Aggregation in Language Models
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Myopic Bayesian Decision Theory for Batch Active Learning with Partial Batch Label Sampling
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
- Gamma Mixture Modeling for Cosine Similarity in Small Language Models
- Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
- P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
- RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
- What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- Robust Federated Inference
- TAP: Two-Stage Adaptive Personalization of Multi-task and Multi-Modal Foundation Models in Federated Learning
- Submodular Context Partitioning and Compression for In-Context Learning
- Vision Function Layer in Multimodal LLMs
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- CURA: Size Isnt All You Need -- A Compact Universal Architecture for On-Device Intelligence
- Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
- Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- Question-Driven Analysis and Synthesis: Building Interpretable Thematic Trees with LLMs for Text Clustering and Controllable Generation
- Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
- SAGE: A Realistic Benchmark for Semantic Understanding
- LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
- Stability of In-Context Learning: A Spectral Coverage Perspective
- Every Character Counts: From Vulnerability to Defense in Phishing Detection
- Efficiently Attacking Memorization Scores
- Enhancing the Effectiveness and Durability of Backdoor Attacks in Federated Learning through Maximizing Task Distinction
- Diversity Boosts AI-Generated Text Detection
- Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
- BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
- AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
- LLM-Guided Co-Training for Text Classification
- Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
- Who to Trust? Aggregating Client Predictions in Federated Distillation
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
- Privacy Preserving In-Context-Learning Framework for Large Language Models
- Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- Label Smoothing++: Enhanced Label Regularization for Training Neural Networks
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- Orthogonal Low-rank Adaptation in Lie Groups for Continual Learning of Large Language Models
- Benchmarking Robust Aggregation in Decentralized Gradient Marketplaces
- CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
- PracMHBench: Re-evaluating Model-Heterogeneous Federated Learning Based on Practical Edge Device Constraints
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
- Guess-and-Learn (G&L): Measuring the Cumulative Error Cost of Cold-Start Adaptation
- Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation
- FFT-MoE: Efficient Federated Fine-Tuning for Foundation Models via Large-scale Sparse MoE under Heterogeneous Edge
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Evaluating Sparse Autoencoders for Monosemantic Representation
- SDEC: Semantic Deep Embedded Clustering
- RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
- Hierarchical Conformal Classification
- SoK: Data Minimization in Machine Learning
- When Explainability Meets Privacy: An Investigation at the Intersection of Post-hoc Explainability and Differential Privacy in the Context of Natural Language Processing
- Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
- Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
- Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
- Integrating attention into explanation frameworks for language and vision transformers
- Gradient Surgery for Safe LLM Fine-Tuning
- SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
- ALScope: A Unified Toolkit for Deep Active Learning
- GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
- DTPA: Dynamic Token-level Prefix Augmentation for Controllable Text Generation
- PromptAL: Sample-Aware Dynamic Soft Prompts for Few-Shot Active Learning
- Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
- VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
- MLP Memory: A Retriever-Pretrained Memory for Large Language Models
- DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
- Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
- A Bayesian Hybrid Parameter-Efficient Fine-Tuning Method for Large Language Models
- Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem
- Modular Delta Merging with Orthogonal Constraints: A Scalable Framework for Continual and Reversible Model Composition
- Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
- GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface
- LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
- Toward Efficient Uncertainty in LLMs through Evidential Knowledge Distillation
- The Tsetlin Machine Goes Deep: Logical Learning and Reasoning With Graphs
- Tiny language models
- GRID: Scalable Task-Agnostic Prompt-Based Continual Learning for Language Models
- VMask: Tunable Label Privacy Protection for Vertical Federated Learning via Layer Masking
- Political Leaning and Politicalness Classification of Texts
- GhostUMAP2: Measuring and Analyzing (r,d)-Stability of UMAP
- FedDPG: An Adaptive Yet Efficient Prompt-tuning Approach in Federated Learning Settings
- Enhancing Cross-task Transfer of Large Language Models via Activation Steering
- DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
- Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
- Adversarial Text Generation with Dynamic Contextual Perturbation
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin
- Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach
- Post-Training Quantization of Generative and Discriminative LSTM Text Classifiers: A Study of Calibration, Class Balance, and Robustness
- Balanced Training Data Augmentation for Aspect-Based Sentiment Analysis
- GNN-CNN: An Efficient Hybrid Model of Convolutional and Graph Neural Networks for Text Representation
- Bridging KAN and MLP: MJKAN, a Hybrid Architecture with Both Efficiency and Expressiveness
- Continual Gradient Low-Rank Projection Fine-Tuning for LLMs
- ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
- Modeling Data Diversity for Joint Instance and Verbalizer Selection in Cold-Start Scenarios
- LADSG: Label-Anonymized Distillation and Similar Gradient Substitution for Label Privacy in Vertical Federated Learning
- Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
- KAIROS: Scalable Model-Agnostic Data Valuation
- Federated Learning-Enabled Hybrid Language Models for Communication-Efficient Token Transmission
- SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
- CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings
- MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification
- Memory Savings at What Cost? A Study of Alternatives to Backpropagation
- GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation
- Little by Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts
- Convolutional Neural Network: Text Classification Model for Open Domain Question Answering System
- Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder
- Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
- FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
- Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs
- SecP-Tuning: Efficient Privacy-Preserving Prompt Tuning for Large Language Models via MPC
- Approximating Language Model Training Data from Weights
- Oldies but Goldies: The Potential of Character N-grams for Romanian Texts
- Automatic Intent-Slot Induction for Dialogue Systems
- Creating User-steerable Projections with Interactive Semantic Mapping
- Theoretically Unmasking Inference Attacks Against LDP-Protected Clients in Federated Vision Models
- UltraSketchLLM: Saliency-Driven Sketching for Ultra-Low Bit LLM Compression
- Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
- Watermarking LLM-Generated Datasets in Downstream Tasks
- Generative or Discriminative? Revisiting Text Classification in the Era of Transformers
- A Fast Unified Model for Parsing and Sentence Understanding
- Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
- The Diffusion Duality
- Residual-PAC Privacy: Automatic Privacy Control Beyond the Gaussian Barrier
- Reliably Bounding False Positives: A Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction
- Graph Neural Networks for Natural Language Processing: A Survey
- TextNAS: A Neural Architecture Search Space tailored for Text Representation
- Clustering and Median Aggregation Improve Differentially Private Inference
- MixUp Training Leads to Reduced Overfitting and Improved Calibration for the Transformer Architecture
- Is linguistically-motivated data augmentation worth it?
- Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
- Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
- HtFLlib: A Comprehensive Heterogeneous Federated Learning Library and Benchmark
- Pruning General Large Language Models into Customized Expert Models
- Robustness in Both Domains: CLIP Needs a Robust Text Encoder
- Esoteric Language Models: A Family of Any-Order Diffusion LLMs
- Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models
- Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
- Adaptive Thresholding for Multi-Label Classification via Global-Local Signal Fusion
- SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
- BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
- Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation
- Domain Pre-training Impact on Representations
- 2kenize: Tying Subword Sequences for Chinese Script Conversion
- Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
- Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
- Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
- Refining Labeling Functions with Limited Labeled Data
- A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods
- Adaptive Budget Allocation for Orthogonal-Subspace Adapter Tuning in LLMs Continual Learning
- The emergent algebraic structure of RNNs and embeddings in NLP
- Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning
- Uncertainty Unveiled: Can Exposure to More In-context Examples Mitigate Uncertainty for Large Language Models?
- Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
- Information-Theoretic Complementary Prompts for Improved Continual Text Classification
- Learning to Select In-Context Demonstration Preferred by Large Language Model
- Learning Extrapolative Sequence Transformations from Markov Chains
- Avoid Forgetting by Preserving Global Knowledge Gradients in Federated Learning with Non-IID Data
- Rethinking the Understanding Ability across LLMs through Mutual Information
- Optimization-Inspired Few-Shot Adaptation for Large Language Models
- ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
- Anchored Diffusion Language Model
- Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
- Beyond Demonstrations: Dynamic Vector Construction from Latent Representations
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
- Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
- Recursive Offloading for LLM Serving in Multi-tier Networks
- When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
- Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures
- Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
- Cost-aware LLM-based Online Dataset Annotation
- Ranking Entity Based on Both of Word Frequency and Word Sematic Features
- Large Language Models as Computable Approximations to Solomonoff Induction
- PRL: Prompts from Reinforcement Learning
- Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
- CtrlDiff: Boosting Large Diffusion Language Models with Dynamic Block Prediction and Controllable Generation
- Transformer-F: A Transformer network with effective methods for learning universal sentence representation
- Deterministic Bounds and Random Estimates of Metric Tensors on Neuromanifolds
- SynDec: A Synthesize-then-Decode Approach for Arbitrary Textual Style Transfer via Large Language Models
- MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuning
- Language Models That Walk the Talk: A Framework for Formal Fairness Certificates
- NeuroGen: Neural Network Parameter Generation via Large Language Models
- Self-Destructive Language Model
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
- The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
- On the Interconnections of Calibration, Quantification, and Classifier Accuracy Prediction under Dataset Shift
- On the Security Risks of ML-based Malware Detection Systems: A Survey
- Source Code Authorship Attribution Does Not Generalize from Competitions to Classrooms
- ZENN: A Thermodynamics-Inspired Computational Framework for Heterogeneous Data-Driven Modeling
- Universal Rules for Fooling Deep Neural Networks based Text Classification
- Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
- No Query, No Access
- Towards Zero-Label Language Learning
- FNBench: Benchmarking Robust Federated Learning against Noisy Labels
- DRIFT: Drift-Resilient Invariant-Feature Transformer for DGA Detection
- The Efficiency of Pre-training with Objective Masking in Pseudo Labeling for Semi-Supervised Text Classification
- Differentiating Emigration from Return Migration of Scholars Using Name-Based Nationality Detection Models
- Emotions in the Loop: A Survey of Affective Computing for Emotional Support
- kFolden: k-Fold Ensemble for Out-Of-Distribution Detection
- Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
- ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
- CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction
- Updated science-wide author databases of standardized citation indicators
- War of Words II: Enriched Models of Law-Making Processes
- Communication-Efficient Wireless Federated Fine-Tuning for Large-Scale AI Models
- Neural information retrieval: at the end of the early years
- Generalized Discrete Diffusion from Snapshots
- 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models
- Learning to Translate from Soft to Hard LLM Prompts
- SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction
- Disinformation: analysis and identification
- A Survey on Parameter-Efficient Fine-Tuning for Foundation Models in Federated Learning
- Hubs and Spokes Learning: Efficient and Scalable Collaborative Machine Learning
- Federated One-Shot Learning with Data Privacy and Objective-Hiding
- Geometry-Aware Localized Watermarking for Copyright Protection in Embedding-as-a-Service
- Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
- TailLoR: Protecting Principal Components in Parameter-Efficient Continual Learning
- MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
- CobwebTM: Probabilistic Concept Formation for Lifelong and Hierarchical Topic Modeling
- RAIR: Retrieval-Augmented Iterative Refinement for Chinese Spelling Correction
- Text Understanding from Scratch
- Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
- Detecting LLM-Generated Text with Performance Guarantees
- FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
- RAGAT-Mind: A Multi-Granular Modeling Approach for Rumor Detection Based on MindSpore
- Anti-adversarial Learning: Desensitizing Prompts for Large Language Models
- The Ultimate Cookbook for Invisible Poison: Crafting Subtle Clean-Label Text Backdoors with Style Attributes
- BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
- DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models
- Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification
- Batch Aggregation: An Approach to Enhance Text Classification with Correlated Augmented Data
- Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
- TextBugger: Generating Adversarial Text Against Real-world Applications
- Manifold-Constrained Sentence Embeddings via Triplet Loss: Projecting Semantics onto Spheres, Tori, and Möbius Strips
- CAPO: Cost-Aware Prompt Optimization
- A comparison of machine learning algorithms for the surveillance of autism spectrum disorder
- Deep Character-Level Click-Through Rate Prediction for Sponsored Search
- Efficient Function Orchestration for Large Language Models
- On Sampling Strategies for Neural Network-based Collaborative Filtering
- D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model
- Benchmark and application of unsupervised classification approaches for univariate data
- "What is relevant in a text document?": An interpretable machine learning approach
- Denoising Multi-Source Weak Supervision for Neural Text Classification
- Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study
- Refining Financial Consumer Complaints through Multi-Scale Model Interaction
- LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
- DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
- Defending Deep Neural Networks against Backdoor Attacks via Module Switching
- Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
- Mixture-of-Personas Language Models for Population Simulation
Discussions
Related