Transformers
GGUF
English
rag
context obedient
TroyDoesAI
Mermaid
Flow
Diagram
Sequence
Map
Context
Accurate
Summarization
Story
Code
Coder
Architecture
Retrieval
Augmented
Generation
AI
LLM
Mistral
LLama
Large Language Model
Retrieval Augmented Generation
Troy Andrew Schultz
LookingForWork
OpenForHire
IdoCoolStuff
Knowledge Graph
Knowledge
Graph
Accelerator
Enthusiast
Chatbot
Personal Assistant
Copilot
lol
tags
Pruned
efficient
smaller
small
local
open
source
open source
quant
quantize
ablated
Ablation
uncensored
unaligned
bad
alignment
Inference Endpoints
imatrix
mradermacher
commited on
Commit
•
7ded845
1
Parent(s):
dc1359b
auto-patch README.md
Browse files
README.md
CHANGED
@@ -1,6 +1,128 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
1 |
<!-- ### quantize_version: 2 -->
|
2 |
<!-- ### output_tensor_quantised: 1 -->
|
3 |
<!-- ### convert_type: hf -->
|
4 |
<!-- ### vocab_type: -->
|
5 |
<!-- ### tags: nicoboss -->
|
6 |
weighted/imatrix quants of https://huggingface.co/TroyDoesAI/Codestral-21B-Pruned
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
1 |
+
---
|
2 |
+
base_model: TroyDoesAI/Codestral-21B-Pruned
|
3 |
+
language:
|
4 |
+
- en
|
5 |
+
library_name: transformers
|
6 |
+
license: apache-2.0
|
7 |
+
quantized_by: mradermacher
|
8 |
+
tags:
|
9 |
+
- rag
|
10 |
+
- context obedient
|
11 |
+
- TroyDoesAI
|
12 |
+
- Mermaid
|
13 |
+
- Flow
|
14 |
+
- Diagram
|
15 |
+
- Sequence
|
16 |
+
- Map
|
17 |
+
- Context
|
18 |
+
- Accurate
|
19 |
+
- Summarization
|
20 |
+
- Story
|
21 |
+
- Code
|
22 |
+
- Coder
|
23 |
+
- Architecture
|
24 |
+
- Retrieval
|
25 |
+
- Augmented
|
26 |
+
- Generation
|
27 |
+
- AI
|
28 |
+
- LLM
|
29 |
+
- Mistral
|
30 |
+
- LLama
|
31 |
+
- Large Language Model
|
32 |
+
- Retrieval Augmented Generation
|
33 |
+
- Troy Andrew Schultz
|
34 |
+
- LookingForWork
|
35 |
+
- OpenForHire
|
36 |
+
- IdoCoolStuff
|
37 |
+
- Knowledge Graph
|
38 |
+
- Knowledge
|
39 |
+
- Graph
|
40 |
+
- Accelerator
|
41 |
+
- Enthusiast
|
42 |
+
- Chatbot
|
43 |
+
- Personal Assistant
|
44 |
+
- Copilot
|
45 |
+
- lol
|
46 |
+
- tags
|
47 |
+
- Pruned
|
48 |
+
- efficient
|
49 |
+
- smaller
|
50 |
+
- small
|
51 |
+
- local
|
52 |
+
- open
|
53 |
+
- source
|
54 |
+
- open source
|
55 |
+
- quant
|
56 |
+
- quantize
|
57 |
+
- ablated
|
58 |
+
- Ablation
|
59 |
+
- 'uncensored '
|
60 |
+
- unaligned
|
61 |
+
- 'bad '
|
62 |
+
- alignment
|
63 |
+
---
|
64 |
+
## About
|
65 |
+
|
66 |
<!-- ### quantize_version: 2 -->
|
67 |
<!-- ### output_tensor_quantised: 1 -->
|
68 |
<!-- ### convert_type: hf -->
|
69 |
<!-- ### vocab_type: -->
|
70 |
<!-- ### tags: nicoboss -->
|
71 |
weighted/imatrix quants of https://huggingface.co/TroyDoesAI/Codestral-21B-Pruned
|
72 |
+
|
73 |
+
<!-- provided-files -->
|
74 |
+
static quants are available at https://huggingface.co/mradermacher/Codestral-21B-Pruned-GGUF
|
75 |
+
## Usage
|
76 |
+
|
77 |
+
If you are unsure how to use GGUF files, refer to one of [TheBloke's
|
78 |
+
READMEs](https://huggingface.co/TheBloke/KafkaLM-70B-German-V0.1-GGUF) for
|
79 |
+
more details, including on how to concatenate multi-part files.
|
80 |
+
|
81 |
+
## Provided Quants
|
82 |
+
|
83 |
+
(sorted by size, not necessarily quality. IQ-quants are often preferable over similar sized non-IQ quants)
|
84 |
+
|
85 |
+
| Link | Type | Size/GB | Notes |
|
86 |
+
|:-----|:-----|--------:|:------|
|
87 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ1_S.gguf) | i1-IQ1_S | 4.8 | for the desperate |
|
88 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ1_M.gguf) | i1-IQ1_M | 5.2 | mostly desperate |
|
89 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ2_XXS.gguf) | i1-IQ2_XXS | 5.9 | |
|
90 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ2_XS.gguf) | i1-IQ2_XS | 6.5 | |
|
91 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ2_S.gguf) | i1-IQ2_S | 6.9 | |
|
92 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ2_M.gguf) | i1-IQ2_M | 7.4 | |
|
93 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q2_K.gguf) | i1-Q2_K | 8.1 | IQ3_XXS probably better |
|
94 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ3_XXS.gguf) | i1-IQ3_XXS | 8.4 | lower quality |
|
95 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ3_XS.gguf) | i1-IQ3_XS | 9.0 | |
|
96 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q3_K_S.gguf) | i1-Q3_K_S | 9.4 | IQ3_XS probably better |
|
97 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ3_S.gguf) | i1-IQ3_S | 9.5 | beats Q3_K* |
|
98 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ3_M.gguf) | i1-IQ3_M | 9.8 | |
|
99 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q3_K_M.gguf) | i1-Q3_K_M | 10.5 | IQ3_S probably better |
|
100 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q3_K_L.gguf) | i1-Q3_K_L | 11.4 | IQ3_M probably better |
|
101 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-IQ4_XS.gguf) | i1-IQ4_XS | 11.6 | |
|
102 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q4_0.gguf) | i1-Q4_0 | 12.3 | fast, low quality |
|
103 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q4_K_S.gguf) | i1-Q4_K_S | 12.3 | optimal size/speed/quality |
|
104 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q4_K_M.gguf) | i1-Q4_K_M | 12.9 | fast, recommended |
|
105 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q5_K_S.gguf) | i1-Q5_K_S | 14.9 | |
|
106 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q5_K_M.gguf) | i1-Q5_K_M | 15.3 | |
|
107 |
+
| [GGUF](https://huggingface.co/mradermacher/Codestral-21B-Pruned-i1-GGUF/resolve/main/Codestral-21B-Pruned.i1-Q6_K.gguf) | i1-Q6_K | 17.7 | practically like static Q6_K |
|
108 |
+
|
109 |
+
Here is a handy graph by ikawrakow comparing some lower-quality quant
|
110 |
+
types (lower is better):
|
111 |
+
|
112 |
+
![image.png](https://www.nethype.de/huggingface_embed/quantpplgraph.png)
|
113 |
+
|
114 |
+
And here are Artefact2's thoughts on the matter:
|
115 |
+
https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9
|
116 |
+
|
117 |
+
## FAQ / Model Request
|
118 |
+
|
119 |
+
See https://huggingface.co/mradermacher/model_requests for some answers to
|
120 |
+
questions you might have and/or if you want some other model quantized.
|
121 |
+
|
122 |
+
## Thanks
|
123 |
+
|
124 |
+
I thank my company, [nethype GmbH](https://www.nethype.de/), for letting
|
125 |
+
me use its servers and providing upgrades to my workstation to enable
|
126 |
+
this work in my free time. Additional thanks to [@nicoboss](https://huggingface.co/nicoboss) for giving me access to his hardware for calculating the imatrix for these quants.
|
127 |
+
|
128 |
+
<!-- end -->
|