vix.ing · top · new · best · stats · spec

Martin Marek

  1. Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
    2025/07/09 by Martin Marek, Marek, Martin, Sanae Lotfi +7 · 4 voices · 12 citations
    Computer Science · Medicine · #Topic Modeling #Natural Language Processing Techniques #Artificial Intelligence in Healthcare and Education