TEXT_Datasets
Datasets for fine-tunning, instruction and evaluation of text models from projecte-aina
- Viewer • Updated • 56.4k • 120 • 22
projecte-aina/ceil
Viewer • Updated • 204k • 9Note Named Entities Recognition
projecte-aina/catalanqa
Viewer • Updated • 21.4k • 300 • 1Note QA dataset
projecte-aina/GuiaCat
Viewer • Updated • 5.75k • 134Note Sentiment analysis
projecte-aina/CaWikiTC
Viewer • Updated • 21k • 17Note Text classification
projecte-aina/ancora-ca-ner
Viewer • Updated • 13.6k • 67 • 2Note Named Entities Recognition
projecte-aina/teca
Viewer • Updated • 21.2k • 167 • 1Note Textual entailment
projecte-aina/viquiquad
Viewer • Updated • 596 • 40Note Extractive-QA
projecte-aina/xquad-ca
Viewer • Updated • 1.19k • 191Note Cross-lingual-QA, Extractive-QA
projecte-aina/WikiCAT_ca
Viewer • Updated • 12.4k • 41Note Text classification
projecte-aina/Parafraseja
Viewer • Updated • 22k • 93Note Paraphrase
projecte-aina/sts-ca
Viewer • Updated • 3.07k • 99 • 1Note Semantic Textual Similarity
projecte-aina/wnli-ca
Viewer • Updated • 852 • 256Note Textual entailmen
projecte-aina/tecla
Viewer • Updated • 113k • 47Note Text classification
projecte-aina/vilaquad
Viewer • Updated • 2.1k • 44Note Extractive-QA
projecte-aina/catalan_general_crawling
Viewer • Updated • 711k • 50Note A 435-million-token web corpus of Catalan mainly intended to pretrain language models and word representations.
projecte-aina/raco_forums
Viewer • Updated • 3.95M • 89 • 2Note A 19-million-sentence corpus of Catalan user-generated text built from the forums mainly intended to pretrain language models and word representations.
projecte-aina/catalan_government_crawling
Viewer • Updated • 71k • 25 • 1Note A 39-million-token web corpus of Catalan mainly intended to pretrain language models and word representations.
projecte-aina/catalan_textual_corpus
Viewer • Updated • 3.06M • 34 • 1Note A 1760-million-token web corpus of Catalan mainly intended to pretrain language models and word representations.
projecte-aina/CoQCat
Viewer • Updated • 6k • 48 • 2Note Conversational QA
projecte-aina/caBreu
Viewer • Updated • 3k • 170Note Summarization
projecte-aina/CaSERa-catalan-stance-emotions-raco
Viewer • Updated • 14k • 29Note Emotion and dynamic stance detection
projecte-aina/InToxiCat
Viewer • Updated • 29.8k • 32 • 1Note Abusive language detection
projecte-aina/UD_Catalan-AnCora
Viewer • Updated • 16.7k • 43 • 1Note POS tagging
projecte-aina/CaSSA-catalan-structured-sentiment-analysis
Viewer • Updated • 6.4k • 17 • 3Note Sentiment analysis
projecte-aina/CaSET-catalan-stance-emotions-twitter
Viewer • Updated • 6.77k • 29 • 2Note Emotion, static stance, and dynamic stance detection.
projecte-aina/COPA-ca
Viewer • Updated • 1k • 215Note Commonsense reasoning
projecte-aina/xnli-ca
Viewer • Updated • 7.5k • 157Note Textual entailment
projecte-aina/casum
Viewer • Updated • 218k • 89Note Summarization
projecte-aina/vilasum
Viewer • Updated • 13.8k • 74Note Summarization
projecte-aina/CATalog
Viewer • Updated • 34.3M • 774 • 7Note Language Modeling
projecte-aina/mgsm_ca
Viewer • Updated • 258 • 227Note Question Answering
projecte-aina/MentorES
Viewer • Updated • 10.2k • 37 • 2Note Instruction Tuning
projecte-aina/MentorCA
Viewer • Updated • 10.2k • 36 • 2Note Instruction Tuning
projecte-aina/openbookqa_ca
Viewer • Updated • 1k • 245Note Question Answering
projecte-aina/PAWS-ca
Viewer • Updated • 53.4k • 193Note Paraphrase Identification
projecte-aina/NLUCat
Updated • 14Note Intent classification, spans identification and examples generation.
projecte-aina/siqa_ca
Viewer • Updated • 1.93k • 1.13kNote Multiple Choice Question Answering
projecte-aina/piqa_ca
Viewer • Updated • 1.84k • 139Note Multiple Choice Question Answering
projecte-aina/xstorycloze_ca
Viewer • Updated • 1.87k • 129Note Multiple Choice Commonsense Reasoning
projecte-aina/arc_ca
Viewer • Updated • 4.42k • 249Note Multiple Choice Question Answering
projecte-aina/oasst1_ca
Viewer • Updated • 5.49k • 15Note Instruction Tuning