What's New
corpus
Description:
This entry includes the first part of the e-book "Krhki jezik / Delicate tongue" by author Ariela Herček (COBISS.SI-ID 275223043; ISBN 978-961-7272-68-0 (ePUB)).
Ariela Herček's collection Delicate Tongue is a bilingual ...
Ta vnos vsebuje 1 datoteko (3.23
MB).
Publicly Available
corpus
Description:
Submission includes the first part of the audiobook "Besedi na sledi" (Following the Word) by author Andrej Blatnik (COBISS.ID: 275429379, ISBN: 978-961-291-541-4).
“Besedi na sledi” is a dynamic, original travelogue—almost ...
Ta vnos vsebuje 3 datotek(e) (99.4
MB).
Publicly Available
corpus
Description:
Submission includes the first part of the audiobook "CAMINO – Poklon Junakom 3. nadstropja" (CAMINO – Gift to the Heroes of the 3rd Floor) by author Anton Krepek (COBISS.ID: 275243779, ISBN: 978-961-291-536-0).
The book ...
Ta vnos vsebuje 3 datotek(e) (135.33
MB).
Publicly Available
Največ ogledov
V preteklem tednu
corpus
Description:
ParlaMint 5.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora ...
Ta vnos vsebuje 31 datotek(e) (5.94
GB).
Publicly Available
corpus
Description:
Janes-Tag is a manually annotated corpus of Slovene Computer-Mediated Communication (CMC). It is meant as a gold-standard training and testing dataset for tokenisation, sentence segmentation, word normalisation, morphosyntactic ...
Ta vnos vsebuje 7 datotek(e) (3.83
MB).
Publicly Available
toolService
Description:
The LIST corpus extraction tool is a Java program for extracting lists from text corpora on the levels of characters, word parts, words, and word sets. It supports VERT and TEI P5 XML formats and outputs .CSV files that ...
Ta vnos vsebuje 1 datoteko (231.07
MB).
Publicly Available