CLARIN Tool Portal

The CLASSLA-Stanza model for lemmatisation of standard Slovenian 2.0

2 resources

This model for lemmatisation of standard Slovenian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the SUK training corpus (http://hdl.handle.net/11356/1747) and using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204) expanded with the MaCoCu-sl Slovene web corpus (http://hdl.handle.net/11356/1517). The estimated F1 of the lemma annotations is ~99.11. The difference to the previous version of the model is that the model was trained using the SUK training corpus and uses new embeddings and the new version of the Slovene morphological lexicon Sloleks 3.0 (http://hdl.handle.net/11356/1745).

Use "The CLASSLA-Stanza model for lemmatisation of standard Slovenian 2.0"

The CLASSLA-Stanza model for lemmatisation of non-standard Croatian 2.1

2 resources

The model for lemmatisation of non-standard Croatian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the hr500k training corpus (http://hdl.handle.net/11356/1792) and the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1793), using the hrLex inflectional lexicon (http://hdl.handle.net/11356/1232). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~94.23. The difference to the previous version of the model is that this version is trained on a combination of two corpora (hr500k, ReLDI-NormTagNER-hr).

Use "The CLASSLA-Stanza model for lemmatisation of non-standard Croatian 2.1"

The CLASSLA-Stanza model for semantic role labeling of standard Slovenian 2.0

2 resources

The model for semantic role labeling of standard Slovenian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the SUK training corpus (http://hdl.handle.net/11356/1747) and using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204) extended with the MaCoCu-sl Slovenian web corpus (http://hdl.handle.net/11356/1517). The estimated F1 of the semantic role annotations is ~76.24. The difference to the previous version of the model is that the model was trained using the SUK training corpus and the updated word embeddings.

Use "The CLASSLA-Stanza model for semantic role labeling of standard Slovenian 2.0"

LVBERT - Latvian BERT

4 resources

LVBERT is the first publicly available monolingual BERT language model pre-trained for Latvian. For training we used the original implementation of BERT on TensorFlow with the whole-word masking and the next sentence prediction objectives. We used BERT-BASE configuration with 12 layers, 768 hidden units, 12 heads, 128 sequence length, 128 mini-batch size and 32,000 token vocabulary.

Use "LVBERT - Latvian BERT"

The CLASSLA-Stanza model for morphosyntactic annotation of non-standard Croatian 2.1

3 resources

This model for morphosyntactic annotation of non-standard Croatian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the hr500k training corpus (http://hdl.handle.net/11356/1792) and the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1793), using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1790). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~92.49. The difference to the previous version of the model is that this version uses the new version of Croatian word embeddings and is trained on a combination of two datasets (hr500k, ReLDI-NormTagNER-hr).

Use "The CLASSLA-Stanza model for morphosyntactic annotation of non-standard Croatian 2.1"

The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Slovenian 1.0

3 resources

This model for morphosyntactic annotation of non-standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and the Janes-Tag corpus (http://hdl.handle.net/11356/1238), using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~96.14.

Use "The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Slovenian 1.0"

The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0

3 resources

This model for morphosyntactic annotation of non-standard Croatian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the hr500k training corpus (http://hdl.handle.net/11356/1210), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1205). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~95.11.

Use "The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0"

The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1

2 resources

The model for lemmatisation of non-standard Serbian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200) and the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1794), using the srLex inflectional lexicon (http://hdl.handle.net/11356/1233). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~94.92. The difference to the previous version of the model is that this version is trained on a combination of two corpora (SETimes.SR, ReLDI-NormTagNER-sr).

Use "The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1"

The CLASSLA-StanfordNLP model for lemmatisation of standard Bulgarian 1.0

2 resources

The model for lemmatisation of standard Bulgarian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the BulTreeBank training corpus (http://hdl.handle.net/11495/D93F-C6E9-65D9-2) and using the Bulgarian inflectional lexicon (Popov, Simov, and Vidinska 1998). The estimated F1 of the lemma annotations is ~98.8.

Use "The CLASSLA-StanfordNLP model for lemmatisation of standard Bulgarian 1.0"

The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.3

3 resources

This model for morphosyntactic annotation of standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~97.06. The difference to the previous version of the model is that the model now also includes the Sloleks inflectional lexicon.

Use "The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.3"

Result filters

Metadata provider

Language

Resource type

Tool task

Availability

Project

Keywords

Active filters:

Search results

The CLASSLA-Stanza model for lemmatisation of standard Slovenian 2.0

The CLASSLA-Stanza model for lemmatisation of non-standard Croatian 2.1

The CLASSLA-Stanza model for semantic role labeling of standard Slovenian 2.0

LVBERT - Latvian BERT

The CLASSLA-Stanza model for morphosyntactic annotation of non-standard Croatian 2.1

The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Slovenian 1.0

The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0

The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1

The CLASSLA-StanfordNLP model for lemmatisation of standard Bulgarian 1.0

The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.3

Result filters

Metadata provider

Language

Resource type

Tool task

Availability

Project

Keywords

Active filters:

Search results

Session recording