RichardErkhov
commited on
uploaded readme
Browse files
README.md
ADDED
@@ -0,0 +1,275 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
1 |
+
Quantization made by Richard Erkhov.
|
2 |
+
|
3 |
+
[Github](https://github.com/RichardErkhov)
|
4 |
+
|
5 |
+
[Discord](https://discord.gg/pvy7H8DZMG)
|
6 |
+
|
7 |
+
[Request more models](https://github.com/RichardErkhov/quant_request)
|
8 |
+
|
9 |
+
|
10 |
+
bubo-bubo-13b - GGUF
|
11 |
+
- Model creator: https://huggingface.co/ibivibiv/
|
12 |
+
- Original model: https://huggingface.co/ibivibiv/bubo-bubo-13b/
|
13 |
+
|
14 |
+
|
15 |
+
| Name | Quant method | Size |
|
16 |
+
| ---- | ---- | ---- |
|
17 |
+
| [bubo-bubo-13b.Q2_K.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q2_K.gguf) | Q2_K | 4.52GB |
|
18 |
+
| [bubo-bubo-13b.IQ3_XS.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.IQ3_XS.gguf) | IQ3_XS | 4.99GB |
|
19 |
+
| [bubo-bubo-13b.IQ3_S.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.IQ3_S.gguf) | IQ3_S | 5.27GB |
|
20 |
+
| [bubo-bubo-13b.Q3_K_S.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q3_K_S.gguf) | Q3_K_S | 5.27GB |
|
21 |
+
| [bubo-bubo-13b.IQ3_M.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.IQ3_M.gguf) | IQ3_M | 5.57GB |
|
22 |
+
| [bubo-bubo-13b.Q3_K.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q3_K.gguf) | Q3_K | 5.9GB |
|
23 |
+
| [bubo-bubo-13b.Q3_K_M.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q3_K_M.gguf) | Q3_K_M | 5.9GB |
|
24 |
+
| [bubo-bubo-13b.Q3_K_L.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q3_K_L.gguf) | Q3_K_L | 6.45GB |
|
25 |
+
| [bubo-bubo-13b.IQ4_XS.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.IQ4_XS.gguf) | IQ4_XS | 6.54GB |
|
26 |
+
| [bubo-bubo-13b.Q4_0.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q4_0.gguf) | Q4_0 | 6.86GB |
|
27 |
+
| [bubo-bubo-13b.IQ4_NL.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.IQ4_NL.gguf) | IQ4_NL | 6.9GB |
|
28 |
+
| [bubo-bubo-13b.Q4_K_S.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q4_K_S.gguf) | Q4_K_S | 6.91GB |
|
29 |
+
| [bubo-bubo-13b.Q4_K.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q4_K.gguf) | Q4_K | 7.33GB |
|
30 |
+
| [bubo-bubo-13b.Q4_K_M.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q4_K_M.gguf) | Q4_K_M | 7.33GB |
|
31 |
+
| [bubo-bubo-13b.Q4_1.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q4_1.gguf) | Q4_1 | 7.61GB |
|
32 |
+
| [bubo-bubo-13b.Q5_0.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q5_0.gguf) | Q5_0 | 8.36GB |
|
33 |
+
| [bubo-bubo-13b.Q5_K_S.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q5_K_S.gguf) | Q5_K_S | 8.36GB |
|
34 |
+
| [bubo-bubo-13b.Q5_K.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q5_K.gguf) | Q5_K | 8.6GB |
|
35 |
+
| [bubo-bubo-13b.Q5_K_M.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q5_K_M.gguf) | Q5_K_M | 8.6GB |
|
36 |
+
| [bubo-bubo-13b.Q5_1.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q5_1.gguf) | Q5_1 | 9.1GB |
|
37 |
+
| [bubo-bubo-13b.Q6_K.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q6_K.gguf) | Q6_K | 9.95GB |
|
38 |
+
| [bubo-bubo-13b.Q8_0.gguf](https://huggingface.co/RichardErkhov/ibivibiv_-_bubo-bubo-13b-gguf/blob/main/bubo-bubo-13b.Q8_0.gguf) | Q8_0 | 12.88GB |
|
39 |
+
|
40 |
+
|
41 |
+
|
42 |
+
|
43 |
+
Original model description:
|
44 |
+
---
|
45 |
+
license: llama2
|
46 |
+
language:
|
47 |
+
- en
|
48 |
+
tags:
|
49 |
+
- summary
|
50 |
+
|
51 |
+
---
|
52 |
+
# Bubo Bubo 13B
|
53 |
+
|
54 |
+
![img](./bubo-bubo.png)
|
55 |
+
|
56 |
+
# Prompting
|
57 |
+
|
58 |
+
## Prompt Template for alpaca style
|
59 |
+
|
60 |
+
```
|
61 |
+
### Instruction:
|
62 |
+
|
63 |
+
<prompt> (without the <>)
|
64 |
+
|
65 |
+
### Response:
|
66 |
+
```
|
67 |
+
|
68 |
+
## Sample Code
|
69 |
+
|
70 |
+
```python
|
71 |
+
import torch
|
72 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
73 |
+
|
74 |
+
torch.set_default_device("cuda")
|
75 |
+
|
76 |
+
model = AutoModelForCausalLM.from_pretrained("ibivibiv/bubo-bubo-13b", torch_dtype="auto", device_config='auto')
|
77 |
+
tokenizer = AutoTokenizer.from_pretrained("ibivibiv/bubo-bubo-13b")
|
78 |
+
|
79 |
+
inputs = tokenizer("### Instruction: Summarize this email chain : <email chain stuff here>.\n### Response:\n", return_tensors="pt", return_attention_mask=False)
|
80 |
+
|
81 |
+
outputs = model.generate(**inputs, max_length=200)
|
82 |
+
text = tokenizer.batch_decode(outputs)[0]
|
83 |
+
print(text)
|
84 |
+
```
|
85 |
+
|
86 |
+
# Model Details
|
87 |
+
* **Trained by**: [ibivibiv](https://huggingface.co/ibivibiv)
|
88 |
+
* **Library**: [HuggingFace Transformers](https://github.com/huggingface/transformers)
|
89 |
+
* **Model type:** **bubo-bubo-13b** is an auto-regressive language model fine tuned on the Llama 2 transformer architecture.
|
90 |
+
* **Language(s)**: English
|
91 |
+
* **Purpose**: Has specific training for summary tasks. This model is targeted towards summarizing communication chains specifically.
|
92 |
+
|
93 |
+
# Benchmark Scores
|
94 |
+
|
95 |
+
I ran the benchmark harness, for curiousity, but this model is completely geared towards summarizing.
|
96 |
+
|
97 |
+
| Test Name | Accuracy |
|
98 |
+
|------------------------------------------------------|----------------------|
|
99 |
+
| all | 0.579149139810157 |
|
100 |
+
| arc:challenge | 0.5631399317406144 |
|
101 |
+
| hellaswag | 0.6317466640111532 |
|
102 |
+
| hendrycksTest-abstract_algebra | 0.32 |
|
103 |
+
| hendrycksTest-anatomy | 0.5481481481481482 |
|
104 |
+
| hendrycksTest-astronomy | 0.5657894736842105 |
|
105 |
+
| hendrycksTest-business_ethics | 0.55 |
|
106 |
+
| hendrycksTest-clinical_knowledge | 0.6 |
|
107 |
+
| hendrycksTest-college_biology | 0.6388888888888888 |
|
108 |
+
| hendrycksTest-college_chemistry | 0.38 |
|
109 |
+
| hendrycksTest-college_computer_science | 0.43 |
|
110 |
+
| hendrycksTest-college_mathematics | 0.34 |
|
111 |
+
| hendrycksTest-college_medicine | 0.5260115606936416 |
|
112 |
+
| hendrycksTest-college_physics | 0.3431372549019608 |
|
113 |
+
| hendrycksTest-computer_security | 0.71 |
|
114 |
+
| hendrycksTest-conceptual_physics | 0.49361702127659574 |
|
115 |
+
| hendrycksTest-econometrics | 0.35964912280701755 |
|
116 |
+
| hendrycksTest-electrical_engineering | 0.5586206896551724 |
|
117 |
+
| hendrycksTest-elementary_mathematics | 0.3439153439153439 |
|
118 |
+
| hendrycksTest-formal_logic | 0.3333333333333333 |
|
119 |
+
| hendrycksTest-global_facts | 0.42 |
|
120 |
+
| hendrycksTest-high_school_biology | 0.6903225806451613 |
|
121 |
+
| hendrycksTest-high_school_chemistry | 0.45320197044334976 |
|
122 |
+
| hendrycksTest-high_school_computer_science | 0.58 |
|
123 |
+
| hendrycksTest-high_school_european_history | 0.6787878787878788 |
|
124 |
+
| hendrycksTest-high_school_geography | 0.7424242424242424 |
|
125 |
+
| hendrycksTest-high_school_government_and_politics | 0.8341968911917098 |
|
126 |
+
| hendrycksTest-high_school_macroeconomics | 0.558974358974359 |
|
127 |
+
| hendrycksTest-high_school_mathematics | 0.3 |
|
128 |
+
| hendrycksTest-high_school_microeconomics | 0.5672268907563025 |
|
129 |
+
| hendrycksTest-high_school_physics | 0.33112582781456956 |
|
130 |
+
| hendrycksTest-high_school_psychology | 0.7577981651376147 |
|
131 |
+
| hendrycksTest-high_school_statistics | 0.4212962962962963 |
|
132 |
+
| hendrycksTest-high_school_us_history | 0.8186274509803921 |
|
133 |
+
| hendrycksTest-high_school_world_history | 0.759493670886076 |
|
134 |
+
| hendrycksTest-human_aging | 0.6547085201793722 |
|
135 |
+
| hendrycksTest-human_sexuality | 0.6412213740458015 |
|
136 |
+
| hendrycksTest-international_law | 0.6776859504132231 |
|
137 |
+
| hendrycksTest-jurisprudence | 0.75 |
|
138 |
+
| hendrycksTest-logical_fallacies | 0.6993865030674846 |
|
139 |
+
| hendrycksTest-machine_learning | 0.41964285714285715 |
|
140 |
+
| hendrycksTest-management | 0.7281553398058253 |
|
141 |
+
| hendrycksTest-marketing | 0.8504273504273504 |
|
142 |
+
| hendrycksTest-medical_genetics | 0.6 |
|
143 |
+
| hendrycksTest-miscellaneous | 0.7624521072796935 |
|
144 |
+
| hendrycksTest-moral_disputes | 0.6560693641618497 |
|
145 |
+
| hendrycksTest-moral_scenarios | 0.4346368715083799 |
|
146 |
+
| hendrycksTest-nutrition | 0.673202614379085 |
|
147 |
+
| hendrycksTest-philosophy | 0.7009646302250804 |
|
148 |
+
| hendrycksTest-prehistory | 0.7067901234567902 |
|
149 |
+
| hendrycksTest-professional_accounting | 0.4645390070921986 |
|
150 |
+
| hendrycksTest-professional_law | 0.45697522816166886 |
|
151 |
+
| hendrycksTest-professional_medicine | 0.5514705882352942 |
|
152 |
+
| hendrycksTest-professional_psychology | 0.6013071895424836 |
|
153 |
+
| hendrycksTest-public_relations | 0.6636363636363637 |
|
154 |
+
| hendrycksTest-security_studies | 0.6448979591836734 |
|
155 |
+
| hendrycksTest-sociology | 0.7611940298507462 |
|
156 |
+
| hendrycksTest-us_foreign_policy | 0.84 |
|
157 |
+
| hendrycksTest-virology | 0.4819277108433735 |
|
158 |
+
| hendrycksTest-world_religions | 0.7894736842105263 |
|
159 |
+
| truthfulqa:mc | 0.4762440289139372 |
|
160 |
+
| winogrande | 0.7616416732438832 |
|
161 |
+
| gsm8k | 0.20621683093252463 |
|
162 |
+
|
163 |
+
|
164 |
+
## Citations
|
165 |
+
|
166 |
+
```
|
167 |
+
@misc{open-llm-leaderboard,
|
168 |
+
author = {Edward Beeching and Clémentine Fourrier and Nathan Habib and Sheon Han and Nathan Lambert and Nazneen Rajani and Omar Sanseviero and Lewis Tunstall and Thomas Wolf},
|
169 |
+
title = {Open LLM Leaderboard},
|
170 |
+
year = {2023},
|
171 |
+
publisher = {Hugging Face},
|
172 |
+
howpublished = "\url{https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard}"
|
173 |
+
}
|
174 |
+
```
|
175 |
+
```
|
176 |
+
@software{eval-harness,
|
177 |
+
author = {Gao, Leo and
|
178 |
+
Tow, Jonathan and
|
179 |
+
Biderman, Stella and
|
180 |
+
Black, Sid and
|
181 |
+
DiPofi, Anthony and
|
182 |
+
Foster, Charles and
|
183 |
+
Golding, Laurence and
|
184 |
+
Hsu, Jeffrey and
|
185 |
+
McDonell, Kyle and
|
186 |
+
Muennighoff, Niklas and
|
187 |
+
Phang, Jason and
|
188 |
+
Reynolds, Laria and
|
189 |
+
Tang, Eric and
|
190 |
+
Thite, Anish and
|
191 |
+
Wang, Ben and
|
192 |
+
Wang, Kevin and
|
193 |
+
Zou, Andy},
|
194 |
+
title = {A framework for few-shot language model evaluation},
|
195 |
+
month = sep,
|
196 |
+
year = 2021,
|
197 |
+
publisher = {Zenodo},
|
198 |
+
version = {v0.0.1},
|
199 |
+
doi = {10.5281/zenodo.5371628},
|
200 |
+
url = {https://doi.org/10.5281/zenodo.5371628}
|
201 |
+
}
|
202 |
+
```
|
203 |
+
```
|
204 |
+
@misc{clark2018think,
|
205 |
+
title={Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge},
|
206 |
+
author={Peter Clark and Isaac Cowhey and Oren Etzioni and Tushar Khot and Ashish Sabharwal and Carissa Schoenick and Oyvind Tafjord},
|
207 |
+
year={2018},
|
208 |
+
eprint={1803.05457},
|
209 |
+
archivePrefix={arXiv},
|
210 |
+
primaryClass={cs.AI}
|
211 |
+
}
|
212 |
+
```
|
213 |
+
```
|
214 |
+
@misc{zellers2019hellaswag,
|
215 |
+
title={HellaSwag: Can a Machine Really Finish Your Sentence?},
|
216 |
+
author={Rowan Zellers and Ari Holtzman and Yonatan Bisk and Ali Farhadi and Yejin Choi},
|
217 |
+
year={2019},
|
218 |
+
eprint={1905.07830},
|
219 |
+
archivePrefix={arXiv},
|
220 |
+
primaryClass={cs.CL}
|
221 |
+
}
|
222 |
+
```
|
223 |
+
```
|
224 |
+
@misc{hendrycks2021measuring,
|
225 |
+
title={Measuring Massive Multitask Language Understanding},
|
226 |
+
author={Dan Hendrycks and Collin Burns and Steven Basart and Andy Zou and Mantas Mazeika and Dawn Song and Jacob Steinhardt},
|
227 |
+
year={2021},
|
228 |
+
eprint={2009.03300},
|
229 |
+
archivePrefix={arXiv},
|
230 |
+
primaryClass={cs.CY}
|
231 |
+
}
|
232 |
+
```
|
233 |
+
```
|
234 |
+
@misc{lin2022truthfulqa,
|
235 |
+
title={TruthfulQA: Measuring How Models Mimic Human Falsehoods},
|
236 |
+
author={Stephanie Lin and Jacob Hilton and Owain Evans},
|
237 |
+
year={2022},
|
238 |
+
eprint={2109.07958},
|
239 |
+
archivePrefix={arXiv},
|
240 |
+
primaryClass={cs.CL}
|
241 |
+
}
|
242 |
+
```
|
243 |
+
```
|
244 |
+
@misc{DBLP:journals/corr/abs-1907-10641,
|
245 |
+
title={{WINOGRANDE:} An Adversarial Winograd Schema Challenge at Scale},
|
246 |
+
author={Keisuke Sakaguchi and Ronan Le Bras and Chandra Bhagavatula and Yejin Choi},
|
247 |
+
year={2019},
|
248 |
+
eprint={1907.10641},
|
249 |
+
archivePrefix={arXiv},
|
250 |
+
primaryClass={cs.CL}
|
251 |
+
}
|
252 |
+
```
|
253 |
+
```
|
254 |
+
@misc{DBLP:journals/corr/abs-2110-14168,
|
255 |
+
title={Training Verifiers to Solve Math Word Problems},
|
256 |
+
author={Karl Cobbe and
|
257 |
+
Vineet Kosaraju and
|
258 |
+
Mohammad Bavarian and
|
259 |
+
Mark Chen and
|
260 |
+
Heewoo Jun and
|
261 |
+
Lukasz Kaiser and
|
262 |
+
Matthias Plappert and
|
263 |
+
Jerry Tworek and
|
264 |
+
Jacob Hilton and
|
265 |
+
Reiichiro Nakano and
|
266 |
+
Christopher Hesse and
|
267 |
+
John Schulman},
|
268 |
+
year={2021},
|
269 |
+
eprint={2110.14168},
|
270 |
+
archivePrefix={arXiv},
|
271 |
+
primaryClass={cs.CL}
|
272 |
+
}
|
273 |
+
```
|
274 |
+
|
275 |
+
|