Transformers
GGUF
English
rag
context obedient
TroyDoesAI
Mermaid
Flow
Diagram
Sequence
Map
Context
Accurate
Summarization
Story
Code
Coder
Architecture
Retrieval
Augmented
Generation
AI
LLM
Mistral
LLama
Large Language Model
Retrieval Augmented Generation
Troy Andrew Schultz
LookingForWork
OpenForHire
IdoCoolStuff
Knowledge Graph
Knowledge
Graph
Accelerator
Enthusiast
Chatbot
Personal Assistant
Copilot
lol
tags
Pruned
efficient
smaller
small
local
open
source
open source
quant
quantize
ablated
Ablation
uncensored
unaligned
bad
alignment
Inference Endpoints
mradermacher
commited on
auto-patch README.md
Browse files
README.md
CHANGED
@@ -1,6 +1,121 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
1 |
<!-- ### quantize_version: 2 -->
|
2 |
<!-- ### output_tensor_quantised: 1 -->
|
3 |
<!-- ### convert_type: hf -->
|
4 |
<!-- ### vocab_type: -->
|
5 |
<!-- ### tags: -->
|
6 |
static quants of https://huggingface.co/TroyDoesAI/Codestral-21B-Pruned
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
1 |
+
---
|
2 |
+
base_model: TroyDoesAI/Codestral-21B-Pruned
|
3 |
+
language:
|
4 |
+
- en
|
5 |
+
library_name: transformers
|
6 |
+
license: apache-2.0
|
7 |
+
quantized_by: mradermacher
|
8 |
+
tags:
|
9 |
+
- rag
|
10 |
+
- context obedient
|
11 |
+
- TroyDoesAI
|
12 |
+
- Mermaid
|
13 |
+
- Flow
|
14 |
+
- Diagram
|
15 |
+
- Sequence
|
16 |
+
- Map
|
17 |
+
- Context
|
18 |
+
- Accurate
|
19 |
+
- Summarization
|
20 |
+
- Story
|
21 |
+
- Code
|
22 |
+
- Coder
|
23 |
+
- Architecture
|
24 |
+
- Retrieval
|
25 |
+
- Augmented
|
26 |
+
- Generation
|
27 |
+
- AI
|
28 |
+
- LLM
|
29 |
+
- Mistral
|
30 |
+
- LLama
|
31 |
+
- Large Language Model
|
32 |
+
- Retrieval Augmented Generation
|
33 |
+
- Troy Andrew Schultz
|
34 |
+
- LookingForWork
|
35 |
+
- OpenForHire
|
36 |
+
- IdoCoolStuff
|
37 |
+
- Knowledge Graph
|
38 |
+
- Knowledge
|
39 |
+
- Graph
|
40 |
+
- Accelerator
|
41 |
+
- Enthusiast
|
42 |
+
- Chatbot
|
43 |
+
- Personal Assistant
|
44 |
+
- Copilot
|
45 |
+
- lol
|
46 |
+
- tags
|
47 |
+
- Pruned
|
48 |
+
- efficient
|
49 |
+
- smaller
|
50 |
+
- small
|
51 |
+
- local
|
52 |
+
- open
|
53 |
+
- source
|
54 |
+
- open source
|
55 |
+
- quant
|
56 |
+
- quantize
|
57 |
+
- ablated
|
58 |
+
- Ablation
|
59 |
+
- 'uncensored '
|
60 |
+
- unaligned
|
61 |
+
- 'bad '
|
62 |
+
- alignment
|
63 |
+
---
|
64 |
+
## About
|
65 |
+
|
66 |
<!-- ### quantize_version: 2 -->
|
67 |
<!-- ### output_tensor_quantised: 1 -->
|
68 |
<!-- ### convert_type: hf -->
|
69 |
<!-- ### vocab_type: -->
|
70 |
<!-- ### tags: -->
|
71 |
static quants of https://huggingface.co/TroyDoesAI/Codestral-21B-Pruned
|
72 |
+
|
73 |
+
<!-- provided-files -->
|
74 |
+
weighted/imatrix quants seem not to be available (by me) at this time. If they do not show up a week or so after the static ones, I have probably not planned for them. Feel free to request them by opening a Community Discussion.
|
75 |
+
## Usage
|
76 |
+
|
77 |
+
If you are unsure how to use GGUF files, refer to one of [TheBloke's
|
78 |
+
READMEs](https://huggingface.co/TheBloke/KafkaLM-70B-German-V0.1-GGUF) for
|
79 |
+
more details, including on how to concatenate multi-part files.
|
80 |
+
|
81 |
+
## Provided Quants
|
82 |
+
|
83 |
+
(sorted by size, not necessarily quality. IQ-quants are often preferable over similar sized non-IQ quants)
|
84 |
+
|
85 |
+
| Link | Type | Size/GB | Notes |
|
86 |
+
|:-----|:-----|--------:|:------|
|
87 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q2_K.gguf) | Q2_K | 8.1 | |
|
88 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.IQ3_XS.gguf) | IQ3_XS | 9.0 | |
|
89 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q3_K_S.gguf) | Q3_K_S | 9.4 | |
|
90 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.IQ3_S.gguf) | IQ3_S | 9.5 | beats Q3_K* |
|
91 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.IQ3_M.gguf) | IQ3_M | 9.8 | |
|
92 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q3_K_M.gguf) | Q3_K_M | 10.5 | lower quality |
|
93 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q3_K_L.gguf) | Q3_K_L | 11.4 | |
|
94 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.IQ4_XS.gguf) | IQ4_XS | 11.7 | |
|
95 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q4_K_S.gguf) | Q4_K_S | 12.3 | fast, recommended |
|
96 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q4_K_M.gguf) | Q4_K_M | 12.9 | fast, recommended |
|
97 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q5_K_S.gguf) | Q5_K_S | 14.9 | |
|
98 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q5_K_M.gguf) | Q5_K_M | 15.3 | |
|
99 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q6_K.gguf) | Q6_K | 17.7 | very good quality |
|
100 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF/resolve/main/Codestral-21B-Pruned.Q8_0.gguf) | Q8_0 | 22.9 | fast, best quality |
|
101 |
+
|
102 |
+
Here is a handy graph by ikawrakow comparing some lower-quality quant
|
103 |
+
types (lower is better):
|
104 |
+
|
105 |
+
![image.png](https://www.nethype.de/huggingface_embed/quantpplgraph.png)
|
106 |
+
|
107 |
+
And here are Artefact2's thoughts on the matter:
|
108 |
+
https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9
|
109 |
+
|
110 |
+
## FAQ / Model Request
|
111 |
+
|
112 |
+
See https://huggingface.co/mradermacher/model_requests for some answers to
|
113 |
+
questions you might have and/or if you want some other model quantized.
|
114 |
+
|
115 |
+
## Thanks
|
116 |
+
|
117 |
+
I thank my company, [nethype GmbH](https://www.nethype.de/), for letting
|
118 |
+
me use its servers and providing upgrades to my workstation to enable
|
119 |
+
this work in my free time.
|
120 |
+
|
121 |
+
<!-- end -->
|