Replication data for: Measure schematicity through information content: A quantitative approach to grammaticalization

DOI

This is a study to propose a quantitative method to compute the schematicity of constructions, which is a key indicator of the level of grammaticalization of morphemes. In this method, to estimate the schematicity of a schema made up of two morphemes, i.e., X_ (X is the target morpheme and _ represents an open slot), we need to know the total token frequency of all types of X_, and the token frequencies of all kinds of elements occurring in the open slot. For example, if we are interested in the schematicity of “_ment”. We need to know the total token frequency of “_ment”, which is the sum of the frequencies of “shipment”, “equipment”, “employment”, “appointment” … (all types of “_ment”). We also need to know the token frequencies of “ship”, “equip”, “employ”, “appoint” … (all types of elements occurring in the open slot). Therefore, the data are morpheme bigrams (2-gram) generated from the English and Chinese corpora showing what morphemes can each morpheme combine with, together with the token frequency of each bigram, and the token frequencies of its two components respectively.

AntConc, 4.2.0

Identifier
DOI https://doi.org/10.18710/APTUHA
Related Identifier IsCitedBy https://doi.org/10.1075/lali.00189.zha
Metadata Access https://dataverse.no/oai?verb=GetRecord&metadataPrefix=oai_datacite&identifier=doi:10.18710/APTUHA
Provenance
Creator Zhang, Liulin ORCID logo
Publisher DataverseNO
Contributor Zhang, Liulin; Soochow University; The Tromsø Repository of Language and Linguistics (TROLLing)
Publication Year 2025
Rights info:eu-repo/semantics/openAccess
OpenAccess true
Contact Zhang, Liulin (Soochow University)
Representation
Resource Type annotated corpus data; Dataset
Format text/plain; text/csv
Size 9134; 4911848; 2222436
Version 1.0
Discipline Humanities