upstage
/

SOLAR-10.7B-Instruct-v1.0

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

killawhale2 commited on Dec 14, 2023

Commit

2b079b2

•

1 Parent(s): 6625d6d

add fine-tuning dataset details

Files changed (1) hide show

README.md +34 -3

README.md CHANGED Viewed

@@ -1,5 +1,12 @@
 ---
 license: apache-2.0
 ---
 # **Meet 10.7B Solar: Elevating Performance with Upstage Depth UP Scaling!**
@@ -19,11 +26,35 @@ Solar 10.7B is an ideal choice for fine-tuning. SOLAR-10.7B offers robustness an
 # **Instruction Fine-Tuning Strategy**
 We utilize state-of-the-art instruction fine-tuning methods including supervised fine-tuning (SFT) and direct preference optimization (DPO) [1].
-Using open source datasets with Alpaca- and OpenOrca-style and generated  synthetic datasets, we apply iterative DPO training, a proprietary alignment strategy, to maximize the performance of our resulting model.
-*Note:* We were careful of data contamination during SFT and DPO, e.g., removing data created using TruthfulQA's prompts.
-[1] Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D. and Finn, C., 2023. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290.
 # **Evaluation Results**

 ---
 license: apache-2.0
+datasets:
+- c-s-ale/alpaca-gpt4-data
+- Open-Orca/OpenOrca
+- Intel/orca_dpo_pairs
+- allenai/ultrafeedback_binarized_cleaned
+language:
+- en
 ---
 # **Meet 10.7B Solar: Elevating Performance with Upstage Depth UP Scaling!**
 # **Instruction Fine-Tuning Strategy**
 We utilize state-of-the-art instruction fine-tuning methods including supervised fine-tuning (SFT) and direct preference optimization (DPO) [1].
+We used a mixture of the following datasets
+- c-s-ale/alpaca-gpt4-data (SFT)
+- Open-Orca/OpenOrca (SFT)
+- in-house generated data utilizing Metamath [2] (SFT, DPO)
+- Intel/orca_dpo_pairs (DPO)
+- allenai/ultrafeedback_binarized_cleaned (DPO)
+where we were careful of data contamination by not using GSM8K samples when generating data and filtering tasks when applicable via the following list.
+```python
+filtering_task_list = [
+    'task228_arc_answer_generation_easy',
+    'ai2_arc/ARC-Challenge:1.0.0',
+    'ai2_arc/ARC-Easy:1.0.0',
+    'task229_arc_answer_generation_hard',
+    'hellaswag:1.1.0',
+    'task1389_hellaswag_completion',
+    'cot_gsm8k',
+    'cot_gsm8k_ii',
+    'drop:2.0.0',
+    'winogrande:1.1.0'
+]
+```
+Using the datasets mentioned above, we apply SFT and iterative DPO training, a proprietary alignment strategy, to maximize the performance of our resulting model.
+[1] Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C.D. and Finn, C., 2023. Direct preference optimization: Your language model is secretly a reward model. NeurIPS.
+[2] Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J.T., Li, Z., Weller, A. and Liu, W., 2023. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284.
 # **Evaluation Results**