-
AI-generated text corpus AI-GenT 1.0
The AI-Generated Text (AI-GenT) corpus is a collection of English and Slovenian texts generated by several large language models. The corpus has been used in comparisons to... -
Slovene Lexicographic QA Fine-Tuning Corpus SloLexQA 1.0
The Slovene Lexicographic QA Fine-Tuning Corpus is a specialized dataset designed to advance the performance of AI models in understanding the structural, grammatical, and... -
Source code and data for the PhD Thesis "On-Premise Medical Information Extra...
Dataset overview This dataset contains source code and annotation guidelines used in the PhD thesis: “On-Premise Medical Information Extraction from German Doctor’s... -
Dataset of annotated collocation-distractor pairs COLLDIST
The dataset contains 59,598 collocation-distractor pairs for 2,856 headwords. Distractor is defined as an incorrect answer/alternative to collocation, which can be similar to... -
Dataset of annotated headword-synonym-distractor triplets SYNDIST
The dataset contains 51,023 headword-synonym-distractor triplets for 5,000 headwords. Distractor is defined as an incorrect answer/alternative to synonym, which can be similar... -
A multilingual benchmark for evaluating metalinguistic knowledge WALS-Bench 1.0
This is a large-scale multilingual benchmark for evaluating metalinguistic knowledge (i.e. explicit knowledge about the structure of languages) in large language models using... -
Slovene instruction-following dataset for large language models GaMS-Instruct...
GaMS-Instruct-MED-Termset is an instruction-following dataset containing 975,060 prompt-response units in Slovene from the medical domain. It focuses on medical terms, with... -
Slovene instruction-following dataset for large language models GaMS-Instruct...
GaMS-Instruct-MED-Anatomy is an instruction-following dataset containing 711,805 prompt-response units in Slovene (with English and Latin terminology). The units form a... -
Slovene instruction-following dataset for large language models GaMS-Instruct...
GaMS-Instruct-PHARMA is an instruction-following dataset designed to fine-tune Slovene large language models to follow instructions in the medical domain, particularly in the... -
Slovene instruction-following dataset for large language models GaMS-Instruct...
GaMS-Instruct-MED is an instruction-following dataset designed to fine-tune Slovene large language models to follow instructions in the medical domain. It consists of units of... -
Slovenian Dataset for Vision-Language Model Instruction-Tuning SLO-VLM-IT-Dat...
This entry contains the SLO-VLM-IT-Dataset, a comprehensive dataset designed for instruction-tuning vision-language models in the Slovenian language. It is composed of five main... -
Slovene instruction-following dataset for large language models GaMS-Instruct...
GaMS-Instruct-MED is an instruction-following dataset designed to fine-tune Slovene large language models to follow instructions in the medical domain. It consists of pairs of... -
Source code and data for the PhD Thesis "Measuring the Contributions of Visio...
This dataset contains source code and data used in the PhD thesis "Measuring the Contributions of Vision and Text Modalities in Multimodal Transformers". The dataset is split... -
Debiasing Algorithm through Model Adaptation
Debiasing Algorithm through Model Adaptation (DAMA) is based on guarding stereotypical gender signals and model editing. DAMA is performed on specific modules prone to convey... -
mlphys101 - Exploring the performance of Large-Language Models in multilingua...
Large-Language Models such as ChatGPT have the potential to revo- lutionize academic teaching in physics in a similar way the electronic calculator, the home computer or the... -
Source code and data for the PhD Thesis "Measuring the Contributions of Visio...
This dataset contains source code and data used in the PhD thesis "Measuring the Contributions of Vision and Text Modalities in Multimodal Transformers". The dataset is split... -
Dataset of Slovene medical texts PoVeJMo-VeMo-Med 1.0
PoVeJMo-VeMo-Med is a dataset containing Slovene medical texts. The bulk of it is comprised of instructions of use for different prescribed drugs. The texts were extracted from... -
Slovene instruction-following dataset for large language models GaMS-Instruct...
GaMS-Instruct-GEN is an instruction-following dataset designed to fine-tune Slovene large language models to follow instructions. It consists of pairs of prompts and responses,... -
Slovene instruction-following dataset for large language models GaMS-Instruct...
GaMS-Instruct-DH is an instruction-following dataset designed to fine-tune Slovene large language models to follow instructions. It consists of pairs of prompts and responses,...
