vix.ing · top · new · best · stats · spec

Nemotron-4 15B Technical Report

2024/02/26 by Jupinder Parmar, Shrimai Prabhumoye, Parmar, Jupinder +51 · 1 voice · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG

paper · pdf · doi:10.48550/arxiv.2402.16819

arxiv published 2024/02/26 · arxiv updated 2024/02/27

Abstract

We introduce Nemotron-4 15B, a 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. Nemotron-4 15B demonstrates strong performance when assessed on English, multilingual, and coding tasks: it outperforms all existing similarly-sized open models on 4 out of 7 downstream evaluation areas and achieves competitive performance to the leading open models in the remaining ones. Specifically, Nemotron-4 15B exhibits the best multilingual capabilities of all similarly-sized models, even outperforming models over four times larger and those explicitly specialized for multilingual tasks.

Cited by

Discussions

Related