SetFit/enron_spam
Viewer • Updated • 33.7k • 6.98k • 21
How to use kauffinger/xlm-roberta-base-finetuned-enron with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-classification", model="kauffinger/xlm-roberta-base-finetuned-enron") # Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("kauffinger/xlm-roberta-base-finetuned-enron")
model = AutoModelForSequenceClassification.from_pretrained("kauffinger/xlm-roberta-base-finetuned-enron", device_map="auto")I trained this model to detect spam in german as there is no german labeled spam mail dataset, and I could not find an already pretrained multilingual model for the enron spam dataset.
Identifying spam mail in any XLM-RoBERTa-supported language. Note that there was no thorough testing on it's intended use - only validation on the enron mail dataset.
Eval on test set of enron spam: