This dataset provides a curated k-mer library (k = 21) built from reference genomes of European endemic diadromous fishes, complemented with outgroup genomes (e.g., human and other non-target species) to help detect and interpret potential contamination in ancient or low-quality DNA samples. The library was generated with KMC (stored as .kmc_pre/.kmc_suf files) and is intended for use with the DeFiS pipeline (Detection of Fauna and flora: Identification of Species), which identifies the most likely species by comparing k-mer dictionaries between reference genomes and sequencing samples, using metrics such as k-mer ratios and Jaccard similarity.
These precomputed reference k-mer sets are designed to speed up analyses and reduce computational cost by allowing repeated comparisons of a sample against multiple candidate reference genomes. The associated metadata table documents the taxonomic coverage, genome sources (GenBank/NCBI accessions), and reference links used to build the library. Information on how these data were generated and how to use them is also available here: https://bios4biol.pages-forge.inrae.fr/DeFiS/