xglm-564M / README.md
VictoriaLinML's picture
Link to fairseq implementation and model card
92e3107
|
raw
history blame
1.08 kB
metadata
license: mit
thumbnail: https://huggingface.co/front/thumbnails/facebook.png
inference: false

XGLM-564M

XGLM-564M is a multilingual autoregressive language model (with 564 million parameters) trained on a balanced corpus of a diverse set of languages totaling 500 billion sub-tokens. It was introduced in the paper Few-shot Learning with Multilingual Language Models by Xi Victoria Lin*, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li* (*Equal Contribution). The original implementation was released in this repository.

Model card

For intended usage of the model, please refer to the model card released by the team releasing XGLM-564M.