calcuis
/

olmo-gguf

Text2Text Generation

Inference Endpoints

Model card Files Files and versions Community

olmo-gguf / README.md

calcuis's picture

Update README.md

20c0a40 verified 1 day ago

|

history blame contribute delete

1.14 kB

	---
	license: apache-2.0
	language:
	- en
	base_model:
	- allenai/OLMo-7B-0724-Instruct-hf
	pipeline_tag: text2text-generation
	datasets:
	- allenai/tulu-v3.1-mix-preview-4096-OLMoE
	---

	## GGUF quantized version of OLMo-7B-0724-Instruct Model

	project original source: [base model](https://huggingface.co/allenai/OLMo-7B-0724-Instruct-hf)

	Q_2_K (not nice)

	Q_3_K_S (acceptable)

	Q_3_K_M is acceptable (good for running with CPU)

	Q_3_K_L (acceptable)

	Q_4_K_S (okay)

	Q_4_K_M is recommanded (balance)

	Q_5_K_S (good)

	Q_5_K_M (good in general)

	Q_6_K is good also; if you want a better result; take this one instead of Q_5_K_M

	Q_8_0 which is very good; need a reasonable size of RAM otherwise you might expect a long wait

	f16 is similar to the original hf model; opt this one or hf also fine; make sure you have a good machine

	### how to run it

	use any connector for interacting with gguf; i.e., [gguf-connector](https://pypi.org/project/gguf-connector/)

	<img src="https://allenai.org/olmo/olmo-7b-animation.gif" alt="OLMo Logo" width="800" style="margin-left:'auto' margin-right:'auto' display:'block'"/>
	this picture is from the base model