baileyarzate
/

whisper-distil-large-v3-atc-english

@@ -17,21 +17,17 @@ tags: []
 This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
-- **Developed by:** [More Information Needed]
-- **Funded by [optional]:** [More Information Needed]
-- **Shared by [optional]:** [More Information Needed]
-- **Model type:** [More Information Needed]
-- **Language(s) (NLP):** [More Information Needed]
 - **License:** [More Information Needed]
-- **Finetuned from model [optional]:** [More Information Needed]
 ### Model Sources [optional]
 <!-- Provide the basic links for the model. -->
-- **Repository:** [More Information Needed]
-- **Paper [optional]:** [More Information Needed]
-- **Demo [optional]:** [More Information Needed]
 ## Uses
@@ -69,46 +65,110 @@ Users (both direct and downstream) should be made aware of the risks, biases and
 ## How to Get Started with the Model
-Use the code below to get started with the model.
-[More Information Needed]
 ## Training Details
 ### Training Data
 <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
 [More Information Needed]
 ### Training Procedure
 <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
 #### Preprocessing [optional]
-[More Information Needed]
 #### Training Hyperparameters
 - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
 #### Speeds, Sizes, Times [optional]
 <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
-[More Information Needed]
 ## Evaluation
 <!-- This section describes the evaluation protocols and provides the results. -->
 ### Testing Data, Factors & Metrics
 #### Testing Data
 <!-- This should link to a Dataset Card if possible. -->
 [More Information Needed]
@@ -122,33 +182,30 @@ Use the code below to get started with the model.
 <!-- These are the evaluation metrics being used, ideally with a description of why. -->
-[More Information Needed]
 ### Results
-[More Information Needed]
 #### Summary
-## Model Examination [optional]
-<!-- Relevant interpretability work for the model goes here -->
-[More Information Needed]
 ## Environmental Impact
 <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
 Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
-- **Hardware Type:** [More Information Needed]
-- **Hours used:** [More Information Needed]
-- **Cloud Provider:** [More Information Needed]
-- **Compute Region:** [More Information Needed]
-- **Carbon Emitted:** [More Information Needed]
 ## Technical Specifications [optional]
@@ -162,15 +219,20 @@ Carbon emissions can be estimated using the [Machine Learning Impact calculator]
 #### Hardware
-[More Information Needed]
 #### Software
-[More Information Needed]
 ## Citation [optional]
 <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
 **BibTeX:**
@@ -180,20 +242,6 @@ Carbon emissions can be estimated using the [Machine Learning Impact calculator]
 [More Information Needed]
-## Glossary [optional]
-<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
-[More Information Needed]
-## More Information [optional]
-[More Information Needed]
-## Model Card Authors [optional]
-[More Information Needed]
 ## Model Card Contact
-[More Information Needed]

 This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
+- **Developed by:** Jesse Arzate
+- **Model type:** Sequence-to-Sequence (Seq2Seq) Transformer-based model
+- **Language(s) (NLP):** English
 - **License:** [More Information Needed]
+- **Finetuned from model [optional]:** Whisper ASR: distil-large-v3
 ### Model Sources [optional]
 <!-- Provide the basic links for the model. -->
+- **Repository:** https://github.com/Vaibhavs10/fast-whisper-finetuning
 ## Uses
 ## How to Get Started with the Model
+Use the code below to get started with the model.
+```python
+from transformers import (
+    AutomaticSpeechRecognitionPipeline,
+    WhisperForConditionalGeneration,
+    WhisperTokenizer,
+    WhisperProcessor,
+)
+from peft import PeftModel, PeftConfig
+peft_model_id = "baileyarzate/whisper-distil-large-v3-atc-english" # huggingface model path
+language = "en"
+task = "transcribe"
+device = 'cuda'
+peft_config = PeftConfig.from_pretrained(peft_model_id)
+model = WhisperForConditionalGeneration.from_pretrained(
+    peft_config.base_model_name_or_path, device_map="cuda"
+).to(device)
+model = PeftModel.from_pretrained(model, peft_model_id).to(device)
+tokenizer = WhisperTokenizer.from_pretrained(peft_config.base_model_name_or_path, language=language, task=task)
+processor = WhisperProcessor.from_pretrained(peft_config.base_model_name_or_path, language=language, task=task)
+feature_extractor = processor.feature_extractor
+forced_decoder_ids = processor.get_decoder_prompt_ids(language=language, task=task)
+pipe = AutomaticSpeechRecognitionPipeline(model=model, tokenizer=tokenizer, feature_extractor=feature_extractor)
+model.config.use_cache = True
+def transcribe(audio):
+    with torch.cuda.amp.autocast():
+        text = pipe(audio, generate_kwargs={"forced_decoder_ids": forced_decoder_ids}, max_new_tokens=255)["text"]
+    return text
+transcriptions_finetuned = []
+for i in tqdm(range(len(df_subset))):
+    # When you only have audio file path
+    #transcriptions_finetuned.append(transcribe(librosa.load(df["path"][i], sr = 16000, offset = df["start"][i], duration = df["stop"][i] - df["start"][i])[0])) #,model
+    # When you have audio array, saves time
+    transcriptions_finetuned.append(transcribe(df_subset['array'].iloc[i]))
+transcriptions_finetuned = pd.DataFrame(transcriptions_finetuned, columns=['transcription_finetuned'])
+df_subset = df_subset.reset_index().drop(columns=['index'])
+df_subset = pd.concat([df_subset, transcriptions_finetuned], axis=1)
+```
 ## Training Details
 ### Training Data
 <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
+Dataset: ATC audio recordings from actual flight operations.
+Size: ~250 hours of annotated data.
 [More Information Needed]
 ### Training Procedure
 <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
+Modeled the procedure after: https://github.com/Vaibhavs10/fast-whisper-finetuning
 #### Preprocessing [optional]
+Preprocessing: Striped leading and trailing whitespaces from transcript sentences. Removed any sentences containing the phrase "UNINTELLIGIBLE" to filter out unclear or garbled speech. Removed filler words such as "ah" or "uh".
 #### Training Hyperparameters
 - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
+```python
+training_args = Seq2SeqTrainingArguments(
+    per_device_train_batch_size=4,
+    gradient_accumulation_steps=2,
+    learning_rate=5e-4,
+    warmup_steps=100,
+    num_train_epochs=3,
+    fp16=True,
+    per_device_eval_batch_size=4,
+    generation_max_length=128,
+    logging_steps=100,
+    save_steps=500,
+    save_total_limit=3,
+    remove_unused_columns=False,  # required as the PeftModel forward doesn't have the signature of the wrapped model's forward
+    label_names=["labels"],  # same reason as above
+)
+```
 #### Speeds, Sizes, Times [optional]
 <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
+Inference time is about 2 samples per second with an RTX A2000.
 ## Evaluation
 <!-- This section describes the evaluation protocols and provides the results. -->
+Final training loss: 0.103
 ### Testing Data, Factors & Metrics
 #### Testing Data
 <!-- This should link to a Dataset Card if possible. -->
+Dataset: ATC audio recordings from actual flight operations.
+Size: ~250 hours of annotated data.
+Randomly sampled 20% of the data with seed = 42.
 [More Information Needed]
 <!-- These are the evaluation metrics being used, ideally with a description of why. -->
+Word Error Rate
+Normalized Word Error Rate
 ### Results
+Mean WER for 500 test samples: 0.145
+  with 95% confidence interval: (0.123, 0.167)
 #### Summary
+[IN PROGRESS]
 ## Environmental Impact
 <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
 Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+- **Hardware Type:** RTX A2000
+- **Hours used:** 24
+- **Cloud Provider:** Private Infrustructure
+- **Compute Region:** Southern California
+- **Carbon Emitted:** 1.57 kg
 ## Technical Specifications [optional]
 #### Hardware
+CPU: AMD EPYC 7313P 16-Core Processor 3.00 GHz
+GPU: NVIDIA RTX A2000
+vRAM: 6GB
+RAM: 128GB
 #### Software
+Windows 11 Enterprise - 21H2
+Python 3.10.14
 ## Citation [optional]
 <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
+[IN PROGRESS]
 **BibTeX:**
 [More Information Needed]
 ## Model Card Contact
+Jesse Arzate: baileyarzate@gmail.com