TOURSKAD is a processed textual dataset designed for topic modeling, dynamic topic modeling, and skill extraction in tourism research. The dataset contains 7,097 preprocessed excerpts derived from scientific publications in the tourism domain, originally retrieved from Scopus and Web of Science. The source corpus covers publications from 1960 to 2024 and was filtered to retain text related to skills, competencies, literacy, and knowledge in tourism.
The deposited files include a training split, a validation split, and a complete source records summary. The training.csv file contains 5,677 rows and the validation.csv file contains 1,420 rows, both with a single text column compatible with OCTIS-based topic modeling workflows. The source_records_summary.csv file contains the complete processed corpus with id, year, split, and text fields, enabling temporal analyses such as dynamic topic modeling.
The dataset was collected on 25 December 2024 and preprocessed between January and April 2025.
The source records were retrieved from Scopus and Web of Science using a two-level search strategy focused on tourism-related publications containing terms associated with skills, competencies, and literacy. The search included English-language research articles, reviews, book chapters, books, and conference proceedings.
The source corpus covers tourism-related scientific publications from 1960 to 2024. Records with fewer than 50 words were excluded. The deposited dataset contains only processed excerpts and minimal temporal metadata; it does not redistribute the original Scopus or Web of Science export files.
The deposited files document the processed corpus used for computational analysis. The original bibliographic records were accessed through Scopus and Web of Science via Universitat Rovira i Virgili institutional access and are not redistributed in full.