HuggingFaceTB/SmolLM-360M continue pretraining on Malaysian context dataset

Continue pretraining on 50B tokens, dataset prepared at https://github.com/malaysia-ai/pretrain-text-dataset/tree/main/smollm

Wandb at https://wandb.ai/huseinzol05/finetune-HuggingFaceTB-SmolLM-360M/

Downloads last month
19
Safetensors
Model size
362M params
Tensor type
BF16
·
Inference API
Unable to determine this model's library. Check the docs .

Model tree for mesolitica/malaysian-HuggingFaceTB-SmolLM-360M

Finetuned
(29)
this model