Replication data for: Measure schematicity through information content: A quantitative approach to grammaticalization

Dataset

DOI

This is a study to propose a quantitative method to compute the schematicity of constructions, which is a key indicator of the level of grammaticalization of morphemes. In this method, to estimate the schematicity of a schema made up of two morphemes, i.e., X_ (X is the target morpheme and _ represents an open slot), we need to know the total token frequency of all types of X_, and the token frequencies of all kinds of elements occurring in the open slot. For example, if we are interested in the schematicity of “_ment”. We need to know the total token frequency of “_ment”, which is the sum of the frequencies of “shipment”, “equipment”, “employment”, “appointment” … (all types of “_ment”). We also need to know the token frequencies of “ship”, “equip”, “employ”, “appoint” … (all types of elements occurring in the open slot). Therefore, the data are morpheme bigrams (2-gram) generated from the English and Chinese corpora showing what morphemes can each morpheme combine with, together with the token frequency of each bigram, and the token frequencies of its two components respectively.

AntConc, 4.2.0

Identifier
DOI	https://doi.org/10.18710/APTUHA
Related Identifier	IsCitedBy https://doi.org/10.1075/lali.00189.zha
Metadata Access	https://dataverse.no/oai?verb=GetRecord&metadataPrefix=oai_datacite&identifier=doi:10.18710/APTUHA

Provenance
Creator	Zhang, Liulin
Publisher	DataverseNO
Contributor	Zhang, Liulin; Soochow University; The Tromsø Repository of Language and Linguistics (TROLLing)
Publication Year	2025
Rights	info:eu-repo/semantics/openAccess
OpenAccess	true
Contact	Zhang, Liulin (Soochow University)

Representation
Resource Type	annotated corpus data; Dataset
Format	text/plain; text/csv
Size	9134; 4911848; 2222436
Version	1.0
Discipline	Humanities