PyTorch model for Slovenian Coreference Resolution

Dataset

PID

Slovenian model for coreference resolution: a neural network based on a customized transformer architecture, usable with the code published on https://github.com/matejklemen/slovene-coreference-resolution. The model is based on the Slovenian CroSloEngual BERT 1.1 model (http://hdl.handle.net/11356/1330). It was trained on the SUK 1.0 training corpus (http://hdl.handle.net/11356/1747), specifically the SentiCoref subcorpus.

Using the evaluation setting where entity mentions are assumed to be correctly pre-detected, the model achieves the following metric values: MUC: precision = 0.931, recall = 0.957, F1 = 0.943 BCubed: precision = 0.887, recall = 0.947, F1 = 0.914 CEAFe: precision = 0.945, recall = 0.893, F1 = 0.916 CoNLL-12: precision = 0.921, recall = 0.932, F1 = 0.924

Identifier
PID	http://hdl.handle.net/11356/1773
Related Identifier	https://doi.org/10.2298/CSIS201120060K
Related Identifier	https://rsdo.slovenscina.eu/en/semantic-resources-and-technologies
Metadata Access	http://www.clarin.si/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:www.clarin.si:11356/1773

Provenance
Creator	Klemen, Matej; Čebular, Martin; Žitnik, Slavko
Publisher	Faculty of Computer and Information Science, University of Ljubljana
Publication Year	2023
Rights	Creative Commons - Attribution 4.0 International (CC BY 4.0); https://creativecommons.org/licenses/by/4.0/; PUB
OpenAccess	true
Contact	info(at)clarin.si

Representation
Language	Slovenian; Slovene
Resource Type	toolService
Format	text/plain; charset=utf-8; application/zip; downloadable_files_count: 1
Discipline	Linguistics