Multi-Dimensional Analysis of Czech

DOI

Original data for a general-purpose multi-dimensional analysis model of register variation in Czech.

This post contains a CSV data set of 137 linguistic features measured on 3428 Czech text chunks, and an R script which performs a factor analysis on this data set. The results of this factor analysis were used as a basis for an 8-dimensional model of register variation in Czech (see Related Publications), following the methodology introduced by Douglas Biber (see e.g. his 1988 seminal work

Variation Across Speech and Writing

for details on the methodology, or his 2014 article

“Using multi-dimensional analysis to explore cross-linguistic universals of register variation”

for a review of MDA results across a variety of languages).

The data is derived from the

Koditex corpus , which aims to be as diversified as possible, covering various forms of spoken and written (both print and on-line) Czech. In compiling this corpus, the purpose was to provide a solid empirical basis for a comprehensive general-purpose model of register variation in Czech.

Apart from this data set and related publications, additional resources pertaining to the project are available via the

czcorpus/mda

GitHub repository.

R: A Language and Environment for Statistical Computing, 3.4.3

psych: Procedures for Personality and Psychological Research (R package), 1.7.8

Identifier
DOI https://doi.org/10.18710/QAJKZW
Related Identifier IsCitedBy https://doi.org/10.1515/cllt-2018-0020
Metadata Access https://dataverse.no/oai?verb=GetRecord&metadataPrefix=oai_datacite&identifier=doi:10.18710/QAJKZW
Provenance
Creator Cvrček, Václav ORCID logo
Publisher DataverseNO
Contributor Lukeš, David; Czech National Corpus; Cvrček, Václav; Komrsková, Zuzana; Poukarová, Petra; Řehořková, Anna; Zasina, Adrian Jan; The Tromsø Repository of Language and Linguistics (TROLLing)
Publication Year 2018
Funding Reference European Regional Development Fund CZ.02.1.01/0.0/0.0/16_013/0001758
Rights CC0 1.0; info:eu-repo/semantics/openAccess; http://creativecommons.org/publicdomain/zero/1.0
OpenAccess true
Contact Lukeš, David (Czech National Corpus)
Representation
Resource Type corpus data; Dataset
Format application/vnd.openxmlformats-officedocument.wordprocessingml.document; application/pdf; text/tab-separated-values; type/x-r-syntax
Size 49288; 627018; 7420627; 1007
Version 1.1
Discipline Humanities
Spatial Coverage Prague, Czech Republic