Dataset description:
The replication data and R code provide documentation for the article "Evaluating dictionaries for metaphor identification".
A total of 2068 tokens were analyzed for metaphor status four times (using four different learner's dictionaries) by two analysts.
The dataset contains 4 files in addition to the ReadMe file:
01_MIPVU_coders_data.txt: containing the decision about metaphorical status made for each word by two coders.
02_MIPVU_interrater_reliability.R: containing the R code to calculate and visualize the inter-rater reliability of the codings documented in File 01.
03_MIPVU_master_dataset.txt: containing the final decisions for metaphorical status of each word, as well as the disagreement scores for each word across the four dictionaries.
04_Dictionary_comparison.R: contains the R code to calculate and visualize each dictionary's disagreement scores, using the data in File 03.
Article abstract:
The 'Macmillan English Dictionary' has long served as the primary reference tool for researchers applying the Metaphor Identification Procedure Vrije Universiteit (MIPVU) for metaphor identification. The dictionary’s sudden online disappearance thus left many with a decision they had not previously faced: Which dictionary can be used in its place? We compare and evaluate four widely used corpus-based online learner’s dictionaries – Cambridge, Collins, Longman, and Oxford – with the aim of determining whether the choice between them affects metaphor identification outcomes. Text samples were taken from the VU Amsterdam Metaphor Corpus covering academic discourse, news, fiction, and conversation. We first assess disagreement scores. Despite minor discrepancies in lexical unit demarcation and metaphor identification – with Collins being the most divergent dictionary and Oxford the one aligning most with the others – overall disagreement was low. Given this high degree of equivalence, we then examine practical aspects of dictionary use. We argue that factors such as ease of documentation and navigation become methodologically relevant and allow researchers to make an informed choice among empirically equivalent dictionaries.
The four dictionaries:
Cambridge Learner’s Dictionary (CAM): See https://dictionary.cambridge.org/dictionary/learner-english/
Collins English Dictionary (COL): See https://www.collinsdictionary.com/dictionary/english
Longman Dictionary of Contemporary English (LM): See https://www.LMonline.com/
Oxford Advanced Learner’s Dictionary (OX): See https://www.oxfordlearnersdictionaries.com/
The analyzed data consists of 500-word text fragments taken from four text registers in the BNC Baby: academic discourse, (transcribed) conversation, fiction and news.
Text files:
academic discourse: [ECV] 39339 words from Feminist perspectives in philosophy. Whitford, M Griffiths, M Macmillan Publishers Ltd Basingstoke 1989 1-109 (File ID ecv-fragment05)
conversation: [KCU] 49751 words from 9 conversations recorded by ‘Julie’ (PS0GF, R 114) between 20 and 22 February 1992 with 6 interlocutors (File ID kcu-fragment02)
fiction: [G0L] 40217 words from The Lucy ghosts. Shah, Eddy Corgi Books London 1993 321-452 (File ID g0l-fragment01)
news: [A80] 10300 words from The Guardian, electronic edition of 1989-11-08: Sport section. Guardian Newspapers Ltd London 1989 (File ID a80-fragment15)
Consult 'Reference Guide to the BNC Baby (second edition) for more information about the texts: http://www.natcorp.ox.ac.uk/corpus/baby/manual.pdf
R, 4.3.1