Add models

Browse files

Signed-off-by: Xin Liu <[email protected]>

Files changed (14) hide show

.gitattributes +1 -0
Qwen1.5-14B-Chat-Q2_K.gguf +3 -0
Qwen1.5-14B-Chat-Q3_K_L.gguf +3 -0
Qwen1.5-14B-Chat-Q3_K_M.gguf +3 -0
Qwen1.5-14B-Chat-Q3_K_S.gguf +3 -0
Qwen1.5-14B-Chat-Q4_0.gguf +3 -0
Qwen1.5-14B-Chat-Q4_K_M.gguf +3 -0
Qwen1.5-14B-Chat-Q4_K_S.gguf +3 -0
Qwen1.5-14B-Chat-Q5_0.gguf +3 -0
Qwen1.5-14B-Chat-Q5_K_M.gguf +3 -0
Qwen1.5-14B-Chat-Q5_K_S.gguf +3 -0
Qwen1.5-14B-Chat-Q6_K.gguf +3 -0
Qwen1.5-14B-Chat-Q8_0.gguf +3 -0
README.md +74 -1

.gitattributes CHANGED Viewed

@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+*.gguf filter=lfs diff=lfs merge=lfs -text

Qwen1.5-14B-Chat-Q2_K.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e99ca36b4a666a2ed4e612771fcf856f623f088b9fd318d9393fd4fcc7581d34
+size 6087290272

Qwen1.5-14B-Chat-Q3_K_L.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:114434dca6e0100a3405b5f27c82e2f693c3e818ad02e8aace39ee9cd1c80699
+size 7840398752

Qwen1.5-14B-Chat-Q3_K_M.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:015b0856a4e8940422e0a68c48e73309e444c404fc5b5fd9cf33a00fafcca0b1
+size 7418264992

Qwen1.5-14B-Chat-Q3_K_S.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5b82377c9d15fa115fb287efdef600b83229ba312d520ad9044198a8effa9689
+size 6949109152

Qwen1.5-14B-Chat-Q4_0.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c918b7621dc26294afaebdbdbe0816d8bb93a56d190a4fe5e19fa9d1fc5f09af
+size 8179322272

Qwen1.5-14B-Chat-Q4_K_M.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:56370835aed67ec30f7f3058bc47188c2a8171792f08c4b19e3c2a2320349a6e
+size 9191034272

Qwen1.5-14B-Chat-Q4_K_S.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:3ef623518e8e15a4be231c346789cc10d909c2810792df8202ae3011d2e3f7aa
+size 8564960672

Qwen1.5-14B-Chat-Q5_0.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:94d51b3cd0d0cfdae161cfd5944c4d816bcfcdd225fa91ecc136e971228097f8
+size 9852784032

Qwen1.5-14B-Chat-Q5_K_M.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:54b4558377aece56f5637425f433ecdbad48e3fe6e374dc41d147bbfb5f56f81
+size 10535996832

Qwen1.5-14B-Chat-Q5_K_S.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:2d00f99648360e50eddb94463d0f52e909d4411920443a1f300e7e7f536be33f
+size 10028092832

Qwen1.5-14B-Chat-Q6_K.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f233bc5be2c732cb637ba7ac52cb9182e7bdca877128ff0eedb1976194641d73
+size 12310158752

Qwen1.5-14B-Chat-Q8_0.gguf ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:cf4f67d157c825014d1f1ff0fa4434539a3160705694e6b3656931166a89232a
+size 15061728672

README.md CHANGED Viewed

@@ -1,3 +1,76 @@
 ---
-license: apache-2.0
 ---

 ---
+base_model: Qwen/Qwen1.5-14B-Chat
+license: other
+license_name: tongyi-qianwen-research
+license_link: >-
+  https://huggingface.co/Qwen/Qwen1.5-14B-Chat/blob/main/LICENSE
+model_creator: Qwen
+model_name: Qwen1.5 14B Chat
+quantized_by: Second State Inc.
+language:
+- en
+pipeline_tag: text-generation
+tags:
+- chat
 ---
+<!-- header start -->
+<!-- 200823 -->
+<div style="width: auto; margin-left: auto; margin-right: auto">
+<img src="https://github.com/second-state/LlamaEdge/raw/dev/assets/logo.svg" style="width: 100%; min-width: 400px; display: block; margin: auto;">
+</div>
+<hr style="margin-top: 1.0em; margin-bottom: 1.0em;">
+<!-- header end -->
+# Qwen1.5-14B-Chat-GGUF
+## Original Model
+[Qwen/Qwen1.5-7B-Chat](https://huggingface.co/Qwen/Qwen1.5-14B-Chat)
+## Run with LlamaEdge
+- LlamaEdge version: [v0.2.15](https://github.com/second-state/LlamaEdge/releases/tag/0.2.15) and above
+- Prompt template
+  - Prompt type: `chatml`
+  - Prompt string
+    ```text
+    <|im_start|>system
+    {system_message}<|im_end|>
+    <|im_start|>user
+    {prompt}<|im_end|>
+    <|im_start|>assistant
+    ```
+- Run as LlamaEdge service
+  ```bash
+  wasmedge --dir .:. --nn-preload default:GGML:AUTO:Qwen1.5-14B-Chat-Q5_K_M.gguf llama-api-server.wasm -p chatml
+  ```
+- Run as LlamaEdge command app
+  ```bash
+  wasmedge --dir .:. --nn-preload default:GGML:AUTO:Qwen1.5-14B-Chat-Q5_K_M.gguf llama-chat.wasm -p chatml
+  ```
+## Quantized GGUF Models
+| Name | Quant method | Bits | Size | Use case |
+| ---- | ---- | ---- | ---- | ----- |
+| [Qwen1.5-7B-Chat-Q2_K.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q2_K.gguf)     | Q2_K   | 2 | 3.10 GB| smallest, significant quality loss - not recommended for most purposes |
+| [Qwen1.5-7B-Chat-Q3_K_L.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q3_K_L.gguf) | Q3_K_L | 3 | 4.22 GB| small, substantial quality loss |
+| [Qwen1.5-7B-Chat-Q3_K_M.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q3_K_M.gguf) | Q3_K_M | 3 | 3.92 GB| very small, high quality loss |
+| [Qwen1.5-7B-Chat-Q3_K_S.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q3_K_S.gguf) | Q3_K_S | 3 | 3.57 GB| very small, high quality loss |
+| [Qwen1.5-7B-Chat-Q4_0.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q4_0.gguf)     | Q4_0   | 4 | 4.51 GB| legacy; small, very high quality loss - prefer using Q3_K_M |
+| [Qwen1.5-7B-Chat-Q4_K_M.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q4_K_M.gguf) | Q4_K_M | 4 | 4.77 GB| medium, balanced quality - recommended |
+| [Qwen1.5-7B-Chat-Q4_K_S.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q4_K_S.gguf) | Q4_K_S | 4 | 4.54 GB| small, greater quality loss |
+| [Qwen1.5-7B-Chat-Q5_0.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q5_0.gguf)     | Q5_0   | 5 | 5.40 GB| legacy; medium, balanced quality - prefer using Q4_K_M |
+| [Qwen1.5-7B-Chat-Q5_K_M.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q5_K_M.gguf) | Q5_K_M | 5 | 5.53 GB| large, very low quality loss - recommended |
+| [Qwen1.5-7B-Chat-Q5_K_S.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q5_K_S.gguf) | Q5_K_S | 5 | 5.4 GB| large, low quality loss - recommended |
+| [Qwen1.5-7B-Chat-Q6_K.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q6_K.gguf)     | Q6_K   | 6 | 6.34 GB| very large, extremely low quality loss |
+| [Qwen1.5-7B-Chat-Q8_0.gguf](https://huggingface.co/second-state/Qwen1.5-7B-Chat-GGUF/blob/main/Qwen1.5-7B-Chat-Q8_0.gguf)     | Q8_0   | 8 | 8.21 GB| very large, extremely low quality loss - not recommended |