Martin Marek
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
2025/07/09 by Martin Marek, Marek, Martin, Sanae Lotfi +7 · 4 voices · 12 citations
Computer Science · Medicine · #Topic Modeling #Natural Language Processing Techniques #Artificial Intelligence in Healthcare and Education