SYN2025: representative corpus of written Czech

PID

Representative corpus of contemporary written (printed) Czech sized 100 MW. It was created as a representation of printed language from 2020–2024 containing a wide range of text types (fiction, professional literature, newspapers etc.). The corpus is lemmatized, morphologically and syntactically annotated by a combination of various methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The vertical format is a sequence of lines. Each of the lines is either a structure (that starts with '') or a token (with a fixed set of tab-separated columns). The columns of the SYN2025 token lines are described in more detail at https://wiki.korpus.cz/doku.php/en:seznamy:syn2025_attributes

The data provided here exactly correspond to those available via the KonText query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) with ordering randomized within the given document.

Identifier
PID http://hdl.handle.net/11234/1-6110
Related Identifier https://wiki.korpus.cz/doku.php/en:cnk:syn2025
Metadata Access http://lindat.mff.cuni.cz/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:lindat.mff.cuni.cz:11234/1-6110
Provenance
Creator Křen, Michal; Cvrček, Václav; Čapka, Tomáš; Hnátková, Milena; Jelínek, Tomáš; Kocek, Jan; Kováříková, Dominika; Křivan, Jan; Marklová, Anna; Petkevič, Vladimír; Skoumalová, Hana; Škrabal, Michal
Publisher Charles University, Faculty of Arts, Department of Linguistics
Publication Year 2025
Rights License Agreement for Czech National Corpus Data; https://lindat.mff.cuni.cz/repository/static/license-cnc-data.html; ACA
OpenAccess true
Contact lindat-help(at)ufal.mff.cuni.cz
Representation
Language Czech
Resource Type corpus
Format text/plain; charset=utf-8; application/x-xz; downloadable_files_count: 1
Discipline Linguistics