Add model

Browse files

Files changed (6) hide show

README.md +92 -0
config.json +96 -0
fairseq/model.pt +3 -0
preprocessor_config.json +9 -0
pytorch_model.bin +3 -0
rinna.png +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,92 @@

+---
+thumbnail: https://github.com/rinnakk/japanese-pretrained-models/blob/master/rinna.png
+language: ja
+license: apache-2.0
+datasets: reazon-research/reazonspeech
+inference: false
+tags:
+  - data2vec
+  - speech
+---
+# `rinna/japanese-data2vec-audio-base`
+![rinna-icon](./rinna.png)
+# Overview
+This is a Japanese data2vec Audio Base model trained by [rinna Co., Ltd.](https://rinna.co.jp/)
+* **Model summary**
+  The model architecture is the same as the [original data2vec Audio Base model](https://huggingface.co/facebook/data2vec-audio-base), which contains 12 transformer layers with 12 attention heads.
+  The model was trained using code from the [official repository](https://github.com/facebookresearch/fairseq/tree/main/examples/data2vec#data2vec), and the detailed training configuration can be found in the same repository and the [original paper](https://ai.meta.com/research/data2vec-a-general-framework-for-self-supervised-learning-in-speech-vision-and-language/).
+* **Training**
+  The model was trained on approximately 19,000 hours of following Japanese speech corpus ReazonSpeech v1.
+  - [ReazonSpeech](https://huggingface.co/datasets/reazon-research/reazonspeech)
+* **Contributors**
+  - [Yukiya Hono](https://huggingface.co/yky-h)
+  - [Kentaro Mitsui](https://huggingface.co/Kentaro321)
+  - [Kei Sawada](https://huggingface.co/keisawada)
+---
+# How to use the model
+```python
+import soundfile as sf
+from transformers import AutoFeatureExtractor, AutoModel
+model_name = "rinna/japanese-data2vec-audio-base"
+feature_extractor = AutoFeatureExtractor.from_pretrained(model_name)
+model = AutoModel.from_pretrained(model_name)
+model.eval()
+raw_speech_16kHz, sr = sf.read(audio_file)
+inputs = feature_extractor(
+    raw_speech_16kHz,
+    return_tensors="pt",
+    sampling_rate=sr,
+)
+outputs = model(**inputs)
+print(f"Input:  {inputs.input_values.size()}")  # [1, #samples]
+print(f"Output: {outputs.last_hidden_state.size()}")  # [1, #frames, 768]
+```
+A fairseq checkpoint file can also be available [here](https://huggingface.co/rinna/japanese-data2vec-audio-base/tree/main/fairseq).
+---
+# How to cite
+```bibtex
+@misc{rinna-japanese-data2vec-audio-base,
+  title={rinna/japanese-data2vec-audio-base},
+  author={Hono, Yukiya and Mitsui, Kentaro and Sawada, Kei},
+  url={https://huggingface.co/rinna/japanese-data2vec-audio-base}
+}
+```
+---
+# Citations
+```bibtex
+@inproceedings{baevski2022data2vec,
+  title={Data2vec: A general framework for self-supervised learning in speech, vision and language},
+  author={Baevski, Alexei and Hsu, Wei-Ning and Xu, Qiantong and Babu, Arun and Gu, Jiatao and Auli, Michael},
+  booktitle={International Conference on Machine Learning},
+  pages={1298--1312},
+  year={2022},
+  organization={PMLR},
+  doi={10.48550/arXiv.2202.03555}
+}
+```
+---
+# License
+[The Apache 2.0 license](https://www.apache.org/licenses/LICENSE-2.0)

config.json ADDED Viewed

	@@ -0,0 +1,96 @@

+{
+  "_name_or_path": "rinna/japanese-data2vec-base",
+  "activation_dropout": 0.1,
+  "adapter_kernel_size": 3,
+  "adapter_stride": 2,
+  "add_adapter": false,
+  "architectures": [
+    "Data2VecAudioModel"
+  ],
+  "attention_dropout": 0.1,
+  "bos_token_id": 1,
+  "classifier_proj_size": 256,
+  "conv_bias": false,
+  "conv_dim": [
+    512,
+    512,
+    512,
+    512,
+    512,
+    512,
+    512
+  ],
+  "conv_kernel": [
+    10,
+    3,
+    3,
+    3,
+    3,
+    2,
+    2
+  ],
+  "conv_pos_kernel_size": 19,
+  "conv_stride": [
+    5,
+    2,
+    2,
+    2,
+    2,
+    2,
+    2
+  ],
+  "ctc_loss_reduction": "sum",
+  "ctc_zero_infinity": false,
+  "eos_token_id": 2,
+  "feat_extract_activation": "gelu",
+  "feat_proj_dropout": 0.0,
+  "final_dropout": 0.1,
+  "hidden_act": "gelu",
+  "hidden_dropout": 0.1,
+  "hidden_size": 768,
+  "initializer_range": 0.02,
+  "intermediate_size": 3072,
+  "layer_norm_eps": 1e-05,
+  "layerdrop": 0.1,
+  "mask_feature_length": 10,
+  "mask_feature_min_masks": 0,
+  "mask_feature_prob": 0.0,
+  "mask_time_length": 10,
+  "mask_time_min_masks": 2,
+  "mask_time_prob": 0.05,
+  "model_type": "data2vec-audio",
+  "num_adapter_layers": 3,
+  "num_attention_heads": 12,
+  "num_conv_pos_embedding_groups": 16,
+  "num_conv_pos_embeddings": 5,
+  "num_feat_extract_layers": 7,
+  "num_hidden_layers": 12,
+  "output_hidden_size": 768,
+  "pad_token_id": 0,
+  "tdnn_dilation": [
+    1,
+    2,
+    3,
+    1,
+    1
+  ],
+  "tdnn_dim": [
+    512,
+    512,
+    512,
+    512,
+    1500
+  ],
+  "tdnn_kernel": [
+    5,
+    3,
+    3,
+    1,
+    1
+  ],
+  "torch_dtype": "float32",
+  "transformers_version": "4.28.1",
+  "use_weighted_layer_sum": false,
+  "vocab_size": 32,
+  "xvector_output_dim": 512
+}

fairseq/model.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8ca2fad0e75704d7d293f0cb2ddcd87603550e0aa5078c4736ea7b0bf4f1d63d
+size 729428441

preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,9 @@

+{
+  "do_normalize": true,
+  "feature_extractor_type": "Wav2Vec2FeatureExtractor",
+  "feature_size": 1,
+  "padding_side": "right",
+  "padding_value": 0.0,
+  "return_attention_mask": true,
+  "sampling_rate": 16000
+}

pytorch_model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4e06d84bb424ca9ca5b11d82348fd04b7c02d60a4a128694c4f4800ecb1ed908
+size 372731429

rinna.png ADDED Viewed