Chizhov, Pavel
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
2025/06/02 by Pierre-Carl Langlais, Pavel Chizhov, Langlais, Pierre-Carl +18 · 8 voices · 7 citations
Social Sciences · #Artificial Intelligence in Law
- BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
2024/09/06 by Chizhov, Pavel, Arnett, Catherine, Korotkova, Elizaveta +1 · 12 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
2025/04/10 by Pavel Chizhov, Chizhov, Pavel, Mattia Nee +5 · 3 voices · 5 citations
#cs.CL
- Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
2025/06/12 by Aleksandra Sorokovikova, Pavel Chizhov, Sorokovikova, Aleksandra +5 · 5 voices · 5 citations
Computer Science · Medicine · Social Sciences · #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI #Topic Modeling #cs.CL
- Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
2025/04/25 by Pierre-Carl Langlais, Pavel Chizhov, Langlais, Pierre-Carl +16 · 1 voice · 1 citation
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Information Retrieval and Search Behavior