Datasets and code of the paper: 'A personal model of trumpery: Linguistic deception detection in a real-world high-stakes setting'

DOI

Paper abstractLanguage use differs between truthful and deceptive statements, but not all differences are consistent across people and contexts, complicating the identification of deceit in individuals. By relying on fact-checked tweets, we show in three studies (Study 1: 469 tweets; Study 2: 484 tweets; Study 3: 24 models) how well personalized linguistic deception detection performs by developing the first deception model tailored to an individual: the 45th US president. First, we found substantial linguistic differences between factually correct and incorrect tweets. We developed a quantitative model and achieved 73% overall accuracy. Second, we tested out-of-sample prediction and achieved 74% overall accuracy. Third, we compared our personalized model to linguistic models previously reported in the literature. Our model outperformed existing models by 5pp, demonstrating the added value of personalized linguistic analysis in real-world settings. Our results indicate that factually incorrect tweets by the US president are not random mistakes of the sender.Additional detailsThe paper is published in Psychological Science (DOI 10.1177/09567976211015941).Datasets and R code are provided.Explanation on how to use the datasets and R code are provided in the methods section of the paper and the supplementary materials. Funded by: European Research Council Starting grant 638408 Bayesian Markets. For more details, see https://https://cordis.europa.eu/project/id/638408For a website with more background information on this paper, please see https://https://apersonalmodeloftrumpery.com/

This entry is a five-file data package totaling 658.9 KB, containing files in .R and .xlsx formats.If you use this dataset, please cite: van der Zee, Sophie; Baillon, Aurelien; Poppe, Ronald; Havrileck, Alice (2021). Datasets and code of the paper: 'A personal model of trumpery: Linguistic deception detection in a real-world high-stakes setting'. Erasmus University Rotterdam (EUR). Dataset. https://doi.org/10.25397/eur.17179514

Identifier
DOI https://doi.org/10.34894/XPCJNJ
Metadata Access https://dataverse.nl/oai?verb=GetRecord&metadataPrefix=oai_datacite&identifier=doi:10.34894/XPCJNJ
Provenance
Creator van der Zee, Sophie; Baillon, Aurelien ORCID logo; Poppe, Ronald; Havrileck, Alice
Publisher DataverseNL
Contributor van der Zee, Sophie
Publication Year 2025
Funding Reference [Grant Label: European Research Council Starting grant 638408 Bayesian Markets; https://cordis.europa.eu/project/id/638408]
Rights CC-BY-4.0; info:eu-repo/semantics/openAccess; http://creativecommons.org/licenses/by/4.0
OpenAccess true
Contact van der Zee, Sophie (Erasmus School of Economics <https://ror.org/057w15z03>)
Representation
Resource Type Dataset
Format application/vnd.openxmlformats-officedocument.spreadsheetml.sheet; type/x-r-syntax
Size 588901; 13087; 13236; 20238; 18747
Version 1.0
Discipline Humanities; Linguistics; Psychology; Social and Behavioural Sciences