vix.ing · top · new · best · stats · spec

Ye, Liang

  1. Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
    2024/12/06 by Qingyuan Li, Bo Zhang, Li, Qingyuan +13 · 7 citations
    Computer Science · #Computational Physics and Python Applications #Topic Modeling