-
Notifications
You must be signed in to change notification settings - Fork 2
Summer 2026 Session 7
Date: Tuesday June 9, 2026 - 16:30–18:00 BST = 17:30–19:00 CEST
Convenors: Valeria Boano (KU Leuven), Francesca Dell'Oro (University of Bologna), Francesco Mambrini (Università Cattolica del Sacro Cuore)
Youtube link: https://www.youtube.com/live/8qRNQPzGRYQ?si=Itg4RBwV6e6dZaFL
Slides: Slides session 7
The first part of this session investigates the potential of semantic annotation and Wikidata linking for enabling broad-scale analysis of a Latin corpus. The focus will be on the annotation of persons in books 2-6 of the Naturalis Historia by Pliny the Elder. In the first part of the presentation, the methodologies used to annotate the text will be outlined, namely Named Entity Recognition and Linking and the subsequent extraction of relevant property values from Wikidata. In the second part, the extracted data regarding the geographical origin and languages spoken by the mentioned individuals will be used to explore a case study concerning Pliny’s attitude towards Roman and Greek persons. Scholarly interpretations of the matter are not univocal, particularly regarding the author’s relationship with the Greeks, with some pointing to a critical stance and others recognizing his great admiration. The exploration of the dataset will shed light on this problem within a distant reading approach.
The second part of this session focuses on anthroponyms within linguistics and philology, based on the recent edition of the famous – and still somewhat mysterious – corpus of lead tablets from Styra (Euboea). Proper names have a linguistic history that runs parallel to that of the appellative lexicon, yet also diverges from it, revealing at various levels connections with the contingent history and culture of a given region. The lead tablets from Styra bear almost exclusively personal names and were used for voting, drawing lots and/or attesting a citizen’s presence. As such, they constitute one of our most important witnesses to democratic practices in Greece at the beginning of the 5th century BCE. It will be outlined how the edition of this peculiar corpus has required a double methodology: a more epigraphical approach for the extant – old and new – tablets, and a more philological one for the lost material. Besides the editorial aspects, the digital edition makes it possible to take into account linguistic aspects of the names, such as their microsyntax, derivational processes, and statistical distribution within Ancient Greek anthroponomastics, thereby shedding light on broader social and cultural dynamics.
The third part focuses on morphosyntactic annotation, introducing treebanks and the workflow for annotating ancient texts semi-automatically. We will start with a definition of treebanks and by introducing the most important international framework for annotation, Universal Annotation (UD). UD provides unified rules and tags to describe the grammatical structure of texts written in multiple languages of the world, ancient as well as modern. Moreover, UD offers a wide variety of models, trained on the available annotation, to analyze texts automatically, and software that can be used to edit and validate the model's output. We will explore the workflow to proceed from a "raw" Greek texts without any annotation to a complete, UD-compliant treebank, then we will use an annotator to inspect and modify the result. Finally, we discuss how we can validate the output to ensure full compliance with the UD standard.
- Fantoli, M., Boano, V., de Graaf, E., Pellizzari di San Girolamo, C.C. (2026). Wikidata as a Knowledge Base for People of the Greco-Roman World. Journal of Open Humanities Data, 12, Art.No. 21. doi: 10.5334/johd.457
- Matilde Garré. LGPN-Ling for the Preservation of Greek Personal Names in a Digital Environment. Digital Classics Online, 2026, 12 (2), pp.124-139. 10.11588/dco.2026.12.112291.
- Serbat, G. (1987). Il y’a Grecs et Grecs! Quel sens donner au prétendu antihéllénisme de Pline? Helmantica, 38, 272–282. url: https://summa.upsa.es/high.raw?id=0000003218&name=00000001.original.pdf.
- The Lexicon of Greek Personal Names (in particular: Naming practices, The formation of names, The ‘meanings’ of names)
- Mambrini, F. (2016). ‘The Ancient Greek Dependency Treebank: Linguistic Annotation in a Teaching Environment’. In Bodard, G., and M. Romanello (edd.), Digital Classics Outside the Echo-Chamber, 83–99. London: Ubiquity Press. https://doi.org/10.5334/bat.f.
- Beersmans, M., de Graaf, E., Van de Cruys, T., & Fantoli, M. (2023). Training and Evaluation of Named Entity Recognition Models for Classical Latin. In A. Anderson, S. Gordin, B. Li, Y. Liu, & M. C. Passarotti (Eds.), Proceedings of the Ancient Language Processing Workshop (pp. 1–12). INCOMA Ltd., Shoumen, Bulgaria. https://aclanthology.org/2023.alp-1.1
- Ehrmann, M., Hamdi, A., Pontes, E. L., Romanello, M., & Doucet, A. (2023). Named Entity Recognition and Classification in Historical Documents: A Survey. ACM Comput. Surv., 56(2), 27:1-27:47. https://doi.org/10.1145/3604931
- Fögen, T. (2013). Scholarship and competitiveness: Pliny the Elder’s attitude towards his predecessors in the Naturalis Historia. In M. Asper (Ed.), Writing Science. Mathematical and Medical Authorship in Ancient Greece (pp. 83–108). De Gruyter. https://doi.org/10.1515/9783110295122.83
- Rydberg-Cox, J. (2021). Modeling the Sources and Topics of Pliny’s Natural History. Umanistica Digitale, (11), 217–229. https://doi.org/10.6092/issn.2532-8816/12521
- Scharpf, P., Breitinger, C., Spitz, A., Meuschke, N., Greiner-Petter, A., Schubotz, M., & Gipp, B. (2026). Entity Linking with Wikidata: A Systematic Literature Review. ACM Comput. Surv., 58(9), 227:1-227:50. https://doi.org/10.1145/3795134
- Wallace-Hadrill, A. (1990). Pliny the Elder and man’s unnatural history. Greece and Rome, 37(1), 80–96. https://doi.org/10.1017/S0017383500029582
- Digital edition of the lead tablets from Styra (Euboea), forthcoming.
- Dell’Oro, F. 2026. Eretria XXVII – Les lamelles de Styra. Nouvelle édition avec étude paléographique, dialectologique et onomastique. Lausanne: ESAG/InFolio
- Dell’Oro, F. 2020. La question de l’authenticité du lot ‘Waddington’ dans le corpus des lamelles de Styra (IG XII 9, 56): l’apport de la linguistique. Historische Sprachforschung 133, 43–61. https://www.jstor.org/stable/27188455?seq=1
- Masson, O. 1992. Les lamelles de plomb de Styra, IG XII 9,56: essai de bilan. Bulletin de Correspondance Hellénique 116, 61–72. https://www.persee.fr/doc/bch_0007-4217_1992_num_116_1_1695
- Marneffe, M.-C. de, Ch. D. Manning, J. Nivre, and D. Zeman (2021). ‘Universal Dependencies’. Computational Linguistics 47, no. 2: 255–308. https://doi.org/10.1162/coli_a_00402
- Fantoli, M., de Graaf, E., Boano, V. I., Pellizzari di San Girolamo, C. C., & Verreth, H. (2025). Replication Data for: Wikidata as a Knowledge Base for People of the Graeco-Roman World (Version V2) [Data set]. Harvard Dataverse. https://doi.org/10.7910/DVN/42QMWG
- LGPN: https://www.lgpn.ox.ac.uk
- LGPN-Ling: https://lgpn-ling.huma-num.fr/about.html
- Universal Dependencies: https://universaldependencies.org
- UDPipe: https://lindat.mff.cuni.cz/services/udpipe/ (or click here)
- Conllueditor: https://github.com/Orange-OpenSource/conllueditor (older version: 2.32.2)
- UD Validation tool: https://github.com/universaldependencies/tools
- Grewmatch (with tutorial): https://universal.grew.fr/?corpus=UD_English-ParTUT@2.18#
Part 1) The exercise consists of a Jupyter Notebook that will guide students through the Python code necessary to process the textual dataset and perform a statistical analysis of the individuals mentioned in the corpus. Students will be provided with a subset of the original corpus, and will be able to reproduce the analysis.
Part 2) Two exercises will be proposed: the first will consist of editing one of the tablets, and the second of providing a linguistic analysis of a proper name. In both cases, the information will be encoded using or slightly adapting the TEI language and schemas.
Part 3) The exercise consists on selecting an Ancient Greek or Latin text, annotate it selecting the appropriate model in UDPipe and edit/validate the annotation using a GUI tool (Conllueditor will be used during the session).