Upload 14 files

Browse files

Files changed (13) hide show

Aira_emissions.csv +2 -0
README.md +156 -0
added_tokens.json +7 -0
config.json +27 -0
generation_config.json +16 -0
lr_scheduler.pt +3 -0
model.safetensors +3 -0
optimizer.pt +3 -0
rng_state.pt +3 -0
special_tokens_map.json +7 -0
tokenizer.json +0 -0
tokenizer.model +3 -0
tokenizer_config.json +80 -0

Aira_emissions.csv ADDED Viewed

	@@ -0,0 +1,2 @@


1	+ timestamp,project_name,run_id,duration,emissions,emissions_rate,cpu_power,gpu_power,ram_power,cpu_energy,gpu_energy,ram_energy,energy_consumed,country_name,country_iso_code,region,cloud_provider,cloud_region,os,python_version,codecarbon_version,cpu_count,cpu_model,gpu_count,gpu_model,longitude,latitude,ram_total_size,tracking_mode,on_cloud,pue
2	+ 2023-12-08T22:42:54,Aira-2,48d59d1a-e7c7-4a88-8d77-f5bfbfe56cb8,31111.21229338646,1.7137292582164891,5.5083975579466225e-05,42.5,0.0,31.305264472961426,0.36728449035220717,2.874421472868672,0.2704877384588287,3.5121937016797027,Singapore,SGP,,,,Linux-5.15.120+-x86_64-with-glibc2.35,3.10.12,2.3.2,12,Intel(R) Xeon(R) CPU @ 2.20GHz,1,1 x NVIDIA A100-SXM4-40GB,103.8503,1.2868,83.48070526123047,machine,N,1.0

README.md ADDED Viewed

	@@ -0,0 +1,156 @@

+---
+license: apache-2.0
+datasets:
+- nicholasKluge/instruct-aira-dataset
+language:
+- en
+metrics:
+- accuracy
+library_name: transformers
+tags:
+- alignment
+- instruction tuned
+- text generation
+- conversation
+- assistant
+pipeline_tag: text-generation
+widget:
+- text: "How should I call you?<|endofinstruction|>"
+  example_title: Greetings
+- text: "Can you explain what is Machine Learning?<|endofinstruction|>"
+  example_title: Machine Learning
+- text: "Do you know anything about virtue ethics?<|endofinstruction|>"
+  example_title: Ethics
+- text: "How can I make my girlfriend happy?<|endofinstruction|>"
+  example_title: Advise
+inference:
+  parameters:
+    repetition_penalty: 1.2
+    temperature: 0.2
+    top_k: 30
+    top_p: 0.3
+    max_new_tokens: 200
+    length_penalty: 0.3
+    early_stopping: true
+co2_eq_emissions:
+  emissions: 1.71
+  source: CodeCarbon
+  training_type: fine-tuning
+  geographical_location: Singapore
+  hardware_used: NVIDIA A100-SXM4-40GB
+---
+# Aira-2-1B1
+`Aira-2` is the second version of the Aira instruction-tuned series. `Aira-2-1B1` is an instruction-tuned model based on [TinyLlama-1.1B](https://huggingface.co/TinyLlama/TinyLlama-1.1B-intermediate-step-955k-token-2T). The model was trained with a dataset composed of prompts and completions generated synthetically by prompting already-tuned models (ChatGPT, Llama, Open-Assistant, etc).
+Check our gradio-demo in [Spaces](https://huggingface.co/spaces/nicholasKluge/Aira-Demo).
+## Details
+- **Size:** 1,261,545,472 parameters
+- **Dataset:** [Instruct-Aira Dataset](https://huggingface.co/datasets/nicholasKluge/instruct-aira-dataset)
+- **Language:** English
+- **Number of Epochs:** 3
+- **Batch size:** 4
+- **Optimizer:** `torch.optim.AdamW` (warmup_steps = 1e2, learning_rate = 5e-4, epsilon = 1e-8)
+- **GPU:** 1 NVIDIA A100-SXM4-40GB
+- **Emissions:** 1.71 KgCO2 (Singapore)
+- **Total Energy Consumption:** 3.51 kWh
+This repository has the [source code](https://github.com/Nkluge-correa/Aira) used to train this model.
+## Usage
+Three special tokens are used to mark the user side of the interaction and the model's response:
+`<|startofinstruction|>`What is a language model?`<|endofinstruction|>`A language model is a probability distribution over a vocabulary.`<|endofcompletion|>`
+```python
+from transformers import AutoTokenizer, AutoModelForCausalLM
+import torch
+device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
+tokenizer = AutoTokenizer.from_pretrained('nicholasKluge/Aira-2-1B1')
+aira = AutoModelForCausalLM.from_pretrained('nicholasKluge/Aira-2-1B1')
+aira.eval()
+aira.to(device)
+question =  input("Enter your question: ")
+inputs = tokenizer(tokenizer.bos_token + question + tokenizer.sep_token,
+  add_special_tokens=False,
+  return_tensors="pt").to(device)
+responses = aira.generate(**inputs,
+	do_sample=True,
+	top_k=50,
+	top_p=0.95,
+	temperature=0.7,
+	num_return_sequences=2)
+print(f"Question: 👤 {question}\n")
+for i, response in  enumerate(responses):
+	print(f'Response {i+1}: 🤖 {tokenizer.decode(response, skip_special_tokens=True).replace(question, "")}')
+```
+The model will output something like:
+```markdown
+>>>Question: 👤 What is the capital of Brazil?
+>>>Response 1: 🤖 The capital of Brazil is Brasília.
+>>>Response 2: 🤖 The capital of Brazil is Brasília.
+```
+## Limitations
+🤥 Generative models can perpetuate the generation of pseudo-informative content, that is, false information that may appear truthful.
+🤬 In certain types of tasks, generative models can produce harmful and discriminatory content inspired by historical stereotypes.
+## Evaluation
+| Model (TinyLlama)                                             | Average   | [ARC](https://arxiv.org/abs/1803.05457) | [TruthfulQA](https://arxiv.org/abs/2109.07958) | [ToxiGen](https://arxiv.org/abs/2203.09509) |
+|---------------------------------------------------------------|-----------|-----------------------------------------|------------------------------------------------|---------------------------------------------|
+| [Aira-2-1B1](https://huggingface.co/nicholasKluge/Aira-2-1B1) | **42.55** | 25.26                                   | **50.81**                                      | **51.59**                                   |
+| TinyLlama/TinyLlama-1.1B-intermediate-step-955k-token-2T      | 37.52     | **30.89**                               | 39.55                                          | 42.13                                       |
+* Evaluations were performed using the [Language Model Evaluation Harness](https://github.com/EleutherAI/lm-evaluation-harness) (by [EleutherAI](https://www.eleuther.ai/)).
+## Cite as 🤗
+```latex
+@misc{nicholas22aira,
+  doi = {10.5281/zenodo.6989727},
+  url = {https://huggingface.co/nicholasKluge/Aira-2-1B1},
+  author = {Nicholas Kluge Corrêa},
+  title = {Aira},
+  year = {2023},
+  publisher = {HuggingFace},
+  journal = {HuggingFace repository},
+}
+```
+## License
+The `Aira-2-1B1` is licensed under the Apache License, Version 2.0. See the [LICENSE](LICENSE) file for more details.
+# [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
+Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_nicholasKluge__Aira-2-1B1)
+| Metric                | Value                     |
+|-----------------------|---------------------------|
+| Avg.                  | 25.19                     |
+| ARC (25-shot)         | 23.21                     |
+| HellaSwag (10-shot)   | 26.97                     |
+| MMLU (5-shot)         | 24.86                     |
+| TruthfulQA (0-shot)   | 50.63                     |
+| Winogrande (5-shot)   | 50.28                     |
+| GSM8K (5-shot)        | 0.0                       |
+| DROP (3-shot)         | 0.39                      |

added_tokens.json ADDED Viewed

	@@ -0,0 +1,7 @@

+{
+  "<|endofcompletion|>": 32001,
+  "<|endofinstruction|>": 32003,
+  "<|pad|>": 32004,
+  "<|startofinstruction|>": 32000,
+  "<|unk|>": 32002
+}

config.json ADDED Viewed

	@@ -0,0 +1,27 @@

+{
+  "_name_or_path": "TinyLlama/TinyLlama-1.1B-intermediate-step-955k-token-2T",
+  "architectures": [
+    "LlamaForCausalLM"
+  ],
+  "attention_bias": false,
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "hidden_act": "silu",
+  "hidden_size": 2048,
+  "initializer_range": 0.02,
+  "intermediate_size": 5632,
+  "max_position_embeddings": 2048,
+  "model_type": "llama",
+  "num_attention_heads": 32,
+  "num_hidden_layers": 22,
+  "num_key_value_heads": 4,
+  "pretraining_tp": 1,
+  "rms_norm_eps": 1e-05,
+  "rope_scaling": null,
+  "rope_theta": 10000.0,
+  "tie_word_embeddings": false,
+  "torch_dtype": "float32",
+  "transformers_version": "4.35.2",
+  "use_cache": false,
+  "vocab_size": 32005
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,16 @@

+{
+  "bos_token_id": 32000,
+  "eos_token_id": 32001,
+  "pad_token_id": 32004,
+  "unk_token_id": 32002,
+  "sep_token_id": 32003,
+  "do_sample": true,
+  "max_new_tokens": 512,
+  "renormalize_logits": true,
+  "repetition_penalty": 1.1,
+  "temperature": 0.3,
+  "top_k": 30,
+  "top_p": 0.3,
+  "transformers_version": "4.35.2",
+  "use_cache": false
+}

lr_scheduler.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4e728a07c8526d9c0607dd05bdf84b8ab489f6a636d7d07fe63d215d167fe8df
+size 1076

model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:297df1581fb78e91cb6945729c2be0c21a02c0c40c81a29b1e359d95d950edf8
+size 4400298456

optimizer.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:1f82f8e8f8468077b36ed645eb8005e1d406dde1af7cac2b948b81979706ad33
+size 8800724018

rng_state.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4c6de1a011acd92b3bf4b92b531ffec0e4cab0f4980f976de55bd76ffd19f487
+size 6246

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,7 @@

+{
+  "bos_token": "<|startofinstruction|>",
+  "eos_token": "<|endofcompletion|>",
+  "pad_token": "<|pad|>",
+  "sep_token": "<|endofinstruction|>",
+  "unk_token": "<|unk|>"
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer.model ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347
+size 499723

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,80 @@

+{
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<unk>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "</s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "32000": {
+      "content": "<|startofinstruction|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "32001": {
+      "content": "<|endofcompletion|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "32002": {
+      "content": "<|unk|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "32003": {
+      "content": "<|endofinstruction|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "32004": {
+      "content": "<|pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|startofinstruction|>",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|endofcompletion|>",
+  "legacy": false,
+  "model_max_length": 1000000000000000019884624838656,
+  "pad_token": "<|pad|>",
+  "padding_side": "right",
+  "sep_token": "<|endofinstruction|>",
+  "sp_model_kwargs": {},
+  "tokenizer_class": "LlamaTokenizer",
+  "unk_token": "<|unk|>",
+  "use_default_system_prompt": false
+}