JamesMW vdmbrsv commited on
Commit
5c64682
·
0 Parent(s):

Duplicate from tabularisai/multilingual-emotion-classification

Browse files

Co-authored-by: Vadim Borisov <vdmbrsv@users.noreply.huggingface.co>

Files changed (6) hide show
  1. .gitattributes +36 -0
  2. README.md +241 -0
  3. config.json +57 -0
  4. model.safetensors +3 -0
  5. tokenizer.json +3 -0
  6. tokenizer_config.json +14 -0
.gitattributes ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,241 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: FacebookAI/xlm-roberta-base
3
+ language:
4
+ - en
5
+ - zh
6
+ - es
7
+ - hi
8
+ - ar
9
+ - bn
10
+ - pt
11
+ - ru
12
+ - ja
13
+ - de
14
+ - id
15
+ - ta
16
+ - vi
17
+ - ko
18
+ - fr
19
+ - tr
20
+ - it
21
+ - pl
22
+ - uk
23
+ - ur
24
+ - nl
25
+ - pa
26
+ - sw
27
+ library_name: transformers
28
+ license: cc-by-nc-4.0
29
+ pipeline_tag: text-classification
30
+ tags:
31
+ - text-classification
32
+ - emotion-classification
33
+ - emotion
34
+ - multi-label-classification
35
+ - synthetic data
36
+ - social-media-analysis
37
+ - customer-feedback
38
+ - product-reviews
39
+ - brand-monitoring
40
+ - multilingual
41
+ - 🇪🇺
42
+ - region:eu
43
+ - synthetic
44
+ ---
45
+
46
+ > [!TIP]
47
+ > 🚀 These models are now available through the Tabularis API.
48
+ > Fast multilingual sentiment + emotion classification in 23 languages with structured outputs and simple pricing.
49
+ >
50
+ > ✅ Free 10K credits/month
51
+ > 📚 Docs + API key: https://tabularis.ai/sentiment-analysis/
52
+
53
+
54
+
55
+ # 🎭 Multilingual Emotion Classification Model (23 Languages, 11 Emotions)
56
+
57
+ [![Join Discord](https://img.shields.io/badge/Discord-Join%20community-5865F2?logo=discord&logoColor=white)](https://discord.gg/sznxwdqBXj)
58
+
59
+ ## Model Details
60
+ - `Model Name:` tabularisai/multilingual-emotion-classification
61
+ - `Base Model:` FacebookAI/xlm-roberta-base
62
+ - `Task:` Multi-label Text Classification (Emotion Recognition)
63
+ - `Languages:` 23 — English, Mandarin Chinese (中文), Spanish (Español), Hindi (हिन्दी), Arabic (العربية), Bengali (বাংলা), Portuguese (Português), Russian (Русский), Japanese (日本語), German (Deutsch), Indonesian (Bahasa Indonesia), Tamil (தமிழ்), Vietnamese (Tiếng Việt), Korean (한국어), French (Français), Turkish (Türkçe), Italian (Italiano), Polish (Polski), Ukrainian (Українська), Urdu (اردو), Dutch (Nederlands), Punjabi (ਪੰਜਾਬੀ), and Swahili.
64
+ - `Number of Classes:` 11 — *anger, contempt, disgust, fear, frustration, gratitude, joy, love, neutral, sadness, surprise*
65
+ - `Label Mode:` Multi-label — each text can be assigned **zero, one, or multiple** emotions (independent sigmoid heads, τ = 0.5).
66
+ - `Usage:`
67
+ - Social media emotion analysis
68
+ - Customer feedback analysis
69
+ - Product review emotion tagging
70
+ - Brand monitoring
71
+ - Conversational AI / chatbot affect tracking
72
+ - Market research
73
+ - Customer service optimization
74
+
75
+ >
76
+
77
+
78
+ ## Model Description
79
+
80
+ This model is a fine-tuned version of `FacebookAI/xlm-roberta-base` for multilingual multi-label emotion classification. It was trained on synthetic multilingual data covering 23 languages and 11 emotion categories, enabling robust emotion detection across languages, registers, and cultural contexts.
81
+
82
+ Unlike single-label sentiment classifiers, this model predicts a **set** of emotions per input — reflecting the reality that utterances often carry mixed affect (e.g. *gratitude + love*, *frustration + sadness*).
83
+
84
+ ### Training Data
85
+
86
+ Trained on synthetic multilingual data generated by advanced LLMs, providing broad coverage of emotion expressions across all 23 supported languages. All labels are multi-hot vectors over the 11 emotion classes.
87
+
88
+ ### Training Procedure
89
+
90
+ - Fine-tuned for 3 epochs with **BCEWithLogitsLoss** (independent-binary-per-label).
91
+ - Cosine LR schedule with 6% warmup, `lr=2e-5`, effective batch size 64.
92
+ - Mixed precision (bf16) on a single A100; max sequence length 192.
93
+ - Per-epoch checkpointing; epoch 3 was selected by validation F1-micro.
94
+
95
+ ## Evaluation (held-out multilingual test set, 11,500 rows)
96
+
97
+ | Metric | Value |
98
+ |---------------------|-------|
99
+ | F1 (micro) | 0.840 |
100
+ | F1 (macro) | 0.839 |
101
+ | Jaccard (samples) | 0.794 |
102
+ | Subset accuracy | 0.640 |
103
+ | Hamming accuracy | 0.953 |
104
+ | AUROC (micro) | 0.980 |
105
+ | Average Precision (micro) | 0.923 |
106
+ | LRAP | 0.936 |
107
+
108
+ Decision threshold: τ = 0.5 applied independently per label.
109
+
110
+ ## Intended Use
111
+
112
+ Ideal for:
113
+ - Multilingual social media emotion monitoring
114
+ - International customer feedback affect analysis
115
+ - Global product review emotion tagging
116
+ - Worldwide brand sentiment & emotion tracking
117
+ - Affect-aware conversational systems
118
+
119
+ ## How to Use
120
+
121
+ ```python
122
+ from transformers import AutoTokenizer, AutoModelForSequenceClassification
123
+ import torch
124
+
125
+ model_name = "tabularisai/multilingual-emotion-classification"
126
+ tokenizer = AutoTokenizer.from_pretrained(model_name)
127
+ model = AutoModelForSequenceClassification.from_pretrained(model_name)
128
+ model.eval()
129
+
130
+ LABELS = ["anger", "contempt", "disgust", "fear", "frustration",
131
+ "gratitude", "joy", "love", "neutral", "sadness", "surprise"]
132
+
133
+ @torch.no_grad()
134
+ def predict_emotions(texts, threshold: float = 0.5):
135
+ inputs = tokenizer(texts, return_tensors="pt", truncation=True,
136
+ padding=True, max_length=192)
137
+ probs = torch.sigmoid(model(**inputs).logits).cpu().numpy()
138
+ results = []
139
+ for row in probs:
140
+ picked = [(LABELS[i], float(row[i])) for i in range(len(LABELS)) if row[i] >= threshold]
141
+ picked.sort(key=lambda x: -x[1])
142
+ results.append(picked or [("neutral", float(row[LABELS.index("neutral")]))])
143
+ return results
144
+
145
+
146
+ texts = [
147
+ # English
148
+ "Thank you so much for helping me, I really appreciate it!",
149
+ "I can't believe they cancelled the flight again, this is ridiculous.",
150
+ # Spanish
151
+ "¡Qué alegría verte después de tanto tiempo!",
152
+ "Estoy muy decepcionado con el servicio.",
153
+ # Chinese
154
+ "收到你的礼物我真的很感动,谢谢你!",
155
+ "这部电影太吓人了,我都不敢一个人看。",
156
+ # Arabic
157
+ "أنا ممتن جدًا لكل ما فعلته من أجلي.",
158
+ "لا أستطيع تحمّل هذا الوضع أكثر من ذلك.",
159
+ # Hindi
160
+ "आपका यह तोहफ़ा देखकर मेरी आँखों में आँसू आ गए।",
161
+ "यह सेवा बिल्कुल घटिया थी, मैं बहुत निराश हूँ।",
162
+ # Japanese
163
+ "久しぶりに会えて本当に嬉しいです!",
164
+ "また電車が遅れた...本当にうんざりする。",
165
+ # French
166
+ "Je suis tellement reconnaissant pour tout ce que tu as fait.",
167
+ "C'est inadmissible, j'en ai assez de cette situation.",
168
+ # Swahili
169
+ "Asante sana kwa msaada wako, nakupenda sana!",
170
+ "Nimechoka kabisa na huduma hii mbaya.",
171
+ ]
172
+
173
+ for t, r in zip(texts, predict_emotions(texts)):
174
+ tags = ", ".join(f"{lbl}({p:.2f})" for lbl, p in r)
175
+ print(f"Text: {t}\nEmotions: {tags}\n")
176
+ ```
177
+
178
+ Using pipelines (returns probability for each of the 11 classes):
179
+
180
+ ```python
181
+ from transformers import pipeline
182
+
183
+ pipe = pipeline(
184
+ "text-classification",
185
+ model="tabularisai/multilingual-emotion-classification",
186
+ function_to_apply="sigmoid",
187
+ top_k=None,
188
+ )
189
+
190
+ print(pipe("I love this product! It's amazing and works perfectly."))
191
+ ```
192
+
193
+ ## Ethical Considerations
194
+
195
+ Synthetic training data reduces annotator bias and broadens language coverage, but real-world validation is strongly advised before deploying in high-stakes settings. Emotion labels are culturally situated — predictions should be treated as probabilistic signals, not ground truth about a person's internal state.
196
+
197
+ ## Citation
198
+
199
+ ```bib
200
+ @misc{borisov2026multilingual,
201
+ title={Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data},
202
+ author={Vadim Borisov},
203
+ year={2026},
204
+ eprint={2604.12633},
205
+ archivePrefix={arXiv},
206
+ primaryClass={cs.CL},
207
+ url={https://arxiv.org/abs/2604.12633},
208
+ }
209
+ ```
210
+
211
+ ## Contact
212
+
213
+ For inquiries, data, private APIs, better models, contact info@tabularis.ai
214
+
215
+ tabularis.ai
216
+
217
+
218
+ <table align="center">
219
+ <tr>
220
+ <td align="center">
221
+ <a href="https://www.linkedin.com/company/tabularis-ai/">
222
+ <img src="https://cdn.jsdelivr.net/gh/simple-icons/simple-icons/icons/linkedin.svg" alt="LinkedIn" width="30" height="30">
223
+ </a>
224
+ </td>
225
+ <td align="center">
226
+ <a href="https://x.com/tabularis_ai">
227
+ <img src="https://cdn.jsdelivr.net/gh/simple-icons/simple-icons/icons/x.svg" alt="X" width="30" height="30">
228
+ </a>
229
+ </td>
230
+ <td align="center">
231
+ <a href="https://github.com/tabularis-ai">
232
+ <img src="https://cdn.jsdelivr.net/gh/simple-icons/simple-icons/icons/github.svg" alt="GitHub" width="30" height="30">
233
+ </a>
234
+ </td>
235
+ <td align="center">
236
+ <a href="https://tabularis.ai">
237
+ <img src="https://cdn.jsdelivr.net/gh/simple-icons/simple-icons/icons/internetarchive.svg" alt="Website" width="30" height="30">
238
+ </a>
239
+ </td>
240
+ </tr>
241
+ </table>
config.json ADDED
@@ -0,0 +1,57 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_cross_attention": false,
3
+ "architectures": [
4
+ "XLMRobertaForSequenceClassification"
5
+ ],
6
+ "attention_probs_dropout_prob": 0.1,
7
+ "bos_token_id": 0,
8
+ "classifier_dropout": null,
9
+ "dtype": "float32",
10
+ "eos_token_id": 2,
11
+ "hidden_act": "gelu",
12
+ "hidden_dropout_prob": 0.1,
13
+ "hidden_size": 768,
14
+ "id2label": {
15
+ "0": "anger",
16
+ "1": "contempt",
17
+ "2": "disgust",
18
+ "3": "fear",
19
+ "4": "frustration",
20
+ "5": "gratitude",
21
+ "6": "joy",
22
+ "7": "love",
23
+ "8": "neutral",
24
+ "9": "sadness",
25
+ "10": "surprise"
26
+ },
27
+ "initializer_range": 0.02,
28
+ "intermediate_size": 3072,
29
+ "is_decoder": false,
30
+ "label2id": {
31
+ "anger": 0,
32
+ "contempt": 1,
33
+ "disgust": 2,
34
+ "fear": 3,
35
+ "frustration": 4,
36
+ "gratitude": 5,
37
+ "joy": 6,
38
+ "love": 7,
39
+ "neutral": 8,
40
+ "sadness": 9,
41
+ "surprise": 10
42
+ },
43
+ "layer_norm_eps": 1e-05,
44
+ "max_position_embeddings": 514,
45
+ "model_type": "xlm-roberta",
46
+ "num_attention_heads": 12,
47
+ "num_hidden_layers": 12,
48
+ "output_past": true,
49
+ "pad_token_id": 1,
50
+ "position_embedding_type": "absolute",
51
+ "problem_type": "multi_label_classification",
52
+ "tie_word_embeddings": true,
53
+ "transformers_version": "5.5.3",
54
+ "type_vocab_size": 1,
55
+ "use_cache": true,
56
+ "vocab_size": 250002
57
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:300fd92943ec7541749bd56fcb3f4c037d2ba3313bf54c0c4046913156fdd3c3
3
+ size 1112232668
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:acbd420e2269cdc1ef45332d3d5c418be4aef6b8cb5a0b7ccae0893485307153
3
+ size 17098086
tokenizer_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": true,
3
+ "backend": "tokenizers",
4
+ "bos_token": "<s>",
5
+ "cls_token": "<s>",
6
+ "eos_token": "</s>",
7
+ "is_local": false,
8
+ "mask_token": "<mask>",
9
+ "model_max_length": 512,
10
+ "pad_token": "<pad>",
11
+ "sep_token": "</s>",
12
+ "tokenizer_class": "XLMRobertaTokenizer",
13
+ "unk_token": "<unk>"
14
+ }