tiiuae
/

falcon-11B

@@ -1,6 +1,6 @@
 # 🚀 Falcon2-11B
-**Falcon2-11B is a 11B parameters causal decoder-only model built by [TII](https://www.tii.ae) and trained over 5,000B tokens of [RefinedWeb](https://huggingface.co/datasets/tiiuae/falcon-refinedweb) enhanced with curated corpora. The model is made available under the Apache 2.0 license.**
 *Paper coming soon 😊.*
@@ -22,7 +22,6 @@ pipeline = transformers.pipeline(
     model=model,
     tokenizer=tokenizer,
     torch_dtype=torch.bfloat16,
-    trust_remote_code=True,
 )
 sequences = pipeline(
    "Can you explain the concepts of Quantum Computing?",
@@ -47,8 +46,8 @@ For fast inference with Falcon, check-out [Text Generation Inference](https://gi
 ### Model Description
-- **Developed by:** [https://www.tii.ae](https://www.tii.ae);
-- **Model type:** Causal decoder-only;
 - **Language(s) (NLP):** English, German, Spanish, French, Italian, Portuguese, Polish, Dutch, Romanian, Czech, Swedish
 - **License:** TII Falcon License 2.0
@@ -109,26 +108,26 @@ for seq in sequences:
 ### Training Data
-Falcon2-11B was trained over 5,000B tokens of [RefinedWeb](https://huggingface.co/datasets/tiiuae/falcon-refinedweb), a high-quality filtered and deduplicated web dataset which we enhanced with curated corpora. It followed a 4 stage training strategy. The first three stages being focused on increasing the context length, from to 2048 to 4096 and finally to 8192 tokens. The last stage aimed to further enhance performance using only high quality data.
-Overall, the data sources included RefinedWeb-English, Refined Web-Europe (en, de, es, fr, it, pt, pl, nl, ro, sv, cs), high quality technical data, code data, and conversational data extracted from public sources.
 The training stages were as follows:
 | **Stage**    | **Context length** | **Tokens** |
 |--------------|-----------------|-------------|
-| Stage 1 | 2048            | 4500B       |
-| Stage 2 | 4096            | 250B        |
-| Stage 3 | 8192            | 250B        |
-| Stage 4 | 8192            | 500B        |
 The data was tokenized with the Falcon-[7B](https://huggingface.co/tiiuae/falcon-7b)/[11B](https://huggingface.co/tiiuae/falcon-11B) tokenizer.
 ### Training Procedure
-Falcon2-11B was trained on 1024 A100 40GB GPUs, using a 3D parallelism strategy (TP=8, PP=1, DP=128) combined with ZeRO and Flash-Attention 2.
 #### Training Hyperparameters
@@ -172,10 +171,8 @@ Falcon2-11B is a causal decoder-only model trained on a causal language modeling
 The architecture is broadly adapted from the GPT-3 paper ([Brown et al., 2020](https://arxiv.org/abs/2005.14165)), with the following differences:
 * **Positionnal embeddings:** rotary ([Su et al., 2021](https://arxiv.org/abs/2104.09864));
-* **Attention:** multiquery ([Shazeer et al., 2019](https://arxiv.org/abs/1911.02150)) and FlashAttention ([Dao et al., 2022](https://arxiv.org/abs/2205.14135));
-* **Decoder-block:** parallel attention/MLP with a two layer norms.
-For multiquery, we are using an internal variant which uses independent key and values per tensor parallel degree.
 | **Hyperparameter** | **Value** | **Comment**                            |
 |--------------------|-----------|----------------------------------------|
@@ -193,7 +190,7 @@ Falcon2-11B was trained on AWS SageMaker, using on average 1024 A100 40GB GPUs i
 #### Software
-Falcon2-11B was trained a custom distributed training codebase, Gigatron. It uses a 3D parallelism approach combined with ZeRO and high-performance Triton kernels (FlashAttention2, etc.)
 ## Citation

 # 🚀 Falcon2-11B
+**Falcon2-11B is a 11B parameters causal decoder-only model built by [TII](https://www.tii.ae) and trained over 5,000B tokens of [RefinedWeb](https://huggingface.co/datasets/tiiuae/falcon-refinedweb) enhanced with curated corpora. The model is made available under the TII Falcon License 2.0, the permissive Apache 2.0-based software license which includes an acceptable use policy that promotes the responsible use of AI.**
 *Paper coming soon 😊.*
     model=model,
     tokenizer=tokenizer,
     torch_dtype=torch.bfloat16,
 )
 sequences = pipeline(
    "Can you explain the concepts of Quantum Computing?",
 ### Model Description
+- **Developed by:** [https://www.tii.ae](https://www.tii.ae)
+- **Model type:** Causal decoder-only
 - **Language(s) (NLP):** English, German, Spanish, French, Italian, Portuguese, Polish, Dutch, Romanian, Czech, Swedish
 - **License:** TII Falcon License 2.0
 ### Training Data
+Falcon2-11B was trained over 5,000B tokens of [RefinedWeb](https://huggingface.co/datasets/tiiuae/falcon-refinedweb), a high-quality filtered and deduplicated web dataset which we enhanced with curated corpora. It followed a four stage training strategy. The first three stages were focused on increasing the context length, from to 2048 to 4096 and finally to 8192 tokens. The last stage aimed to further enhance performance using only high quality data.
+Overall, the data sources included RefinedWeb-English, Refined Web-Europe (cs, de, es, fr, it, nl, pl, pt, ro, sv), high quality technical data, code data, and conversational data extracted from public sources.
 The training stages were as follows:
 | **Stage**    | **Context length** | **Tokens** |
 |--------------|-----------------|-------------|
+| Stage 1 | 2048            | 4500 B       |
+| Stage 2 | 4096            | 250 B        |
+| Stage 3 | 8192            | 250 B        |
+| Stage 4 | 8192            | 500 B        |
 The data was tokenized with the Falcon-[7B](https://huggingface.co/tiiuae/falcon-7b)/[11B](https://huggingface.co/tiiuae/falcon-11B) tokenizer.
 ### Training Procedure
+Falcon2-11B was trained on 1024 A100 40GB GPUs for the majority of the training, using a 3D parallelism strategy (TP=8, PP=1, DP=128) combined with ZeRO and Flash-Attention 2.
 #### Training Hyperparameters
 The architecture is broadly adapted from the GPT-3 paper ([Brown et al., 2020](https://arxiv.org/abs/2005.14165)), with the following differences:
 * **Positionnal embeddings:** rotary ([Su et al., 2021](https://arxiv.org/abs/2104.09864));
+* **Attention:** multiquery ([Shazeer et al., 2019](https://arxiv.org/abs/1911.02150)) and FlashAttention-2 ([Dao, 2023](https://arxiv.org/abs/2307.08691));
+* **Decoder-block:** parallel attention/MLP.
 | **Hyperparameter** | **Value** | **Comment**                            |
 |--------------------|-----------|----------------------------------------|
 #### Software
+Falcon2-11B was trained a custom distributed training codebase, Gigatron. It uses a 3D parallelism approach combined with ZeRO, high-performance Triton kernels and FlashAttention-2.
 ## Citation