Manually annotated corpora

Item

This is a list of manually annotated corpora that are available as part of the CLARIN Resource Families initiative.

Manual corpora are collections of texts containing manually validated or manually assigned linguistic information, such as morphosyntactic tags, lemmas, syntactic parses, named entities etc. These corpora can be used to train new language annotation tools as well as to test the accuracy of existing annotation tools. 

The corpora and corpus collections are classified into 6 categories based on the type of manual annotation:

Format
Language
License
Created
Last updated
Source of item
Collections
Training Discovery Toolkit