AI, Language Technology, and the Past: A Week at Athena Research Center’s Summer School
A Week of Intense Learning in Athens, Greece: My Report on Attending the Summer School in Language Technology and Artificial Intelligence for Digital Humanities, organized by the Institute for Language Speech Processing, Athena Research Center
Last month, from 21 to 25 September, I had the privilege of being one of the 14 students who were accepted to attend the Summer School in Language Technology and Artificial Intelligence for Digital Humanities, organized by the CLARIN: EL research infrastructure and hosted by the Institute for Language Speech Processing of the Athena Research Center, and also sponsored by ATRIUM TNA. The location and venue were perfect and motivated me to gain new knowledge.
Image 1. Acropolis of Athens.
This five-day summer school focused on Language Technology (LT) and Artificial Intelligence (AI) for Digital Humanities (DH). In short, we were introduced to the digital methods for the humanities and taught how LT and AI technologies can be applied to the framework of digital humanities, as well as to our own fields ( source ). This summer school is specifically designed for graduate and postgraduate students, researchers, and professionals who want to deepen their knowledge of modern AI applications to DH. The lecturers from different Greek research institutions and universities guided us through the week with lectures and practical hands-on sessions.
Image 2. A picture of me at the Summer School. Photo credit: Iro Tsiouli.
Day 1, Monday 21 September
This was the first day of the summer school, when I met my 14 colleagues from Albania, Brazil, France, Germany, Kosovo, Poland, Portugal, Slovakia, Spain, Switzerland, and the UK.
The first lecture focused on CLARIN: EL in the age of LLMs: how the old can feed the new with Kanella Pouli. In this lecture, I learned about the main timelines of modern Large Language Models (LLMs), as well as the major research on ChatGPT. I also learned more in depth about what CLARIN is, as well as about its internal structure. Essentially, CLARIN is a European platform designed to provide easy and accessible access to language technologies, created for researchers in Digital Humanities and Social Sciences. Depending on the country, each center has its own “level”. For example, in Portugal there are only two centers, B and K. The B center is the backbone of CLARIN and includes repositories and research centers hosted by universities, while K-centers are knowledge centers providing specialized information. After learning the theoretical part, we moved to the hands-on session, where we were shown how to extract data from software using LLMs.
The second lecture focused on voice note-taking for Archaeological Fieldwork and was delivered by Chara Tsoukala. Despite not being my area, I found it interesting how note taking in done in archaeological fieldwork, which is quite similar to the linguistic fieldwork.
The third and last lecture of the day was focused on linguistic annotation, delivered by Professor Prokopis Prokopidis. I learned about the foundations and methodology of linguistic annotation, use cases, and hybrid annotation with a human in the loop using modern AI systems. We finished the lecture with a hands-on session in Google Colab.
Day 2, Tuesday 22 September
The second day was fully dedicated to AI and LLMs. We started with an introduction to AI and Machine Learning with Dimitris Galantis and Vassilis Katsouros. After being introduced to the main historical timeline of AI, whose roots go deep into the 1950s of the last century, we moved to the more in-depth study of Machine Learning (ML), which is a sub-field of AI concerned with the development and study of algorithms that can learn from data and generalise to unseen data, as well as perform tasks without being explicitly programmed.
In the second lecture, regarding the foundations of LLMs, we acquired more in-depth knowledge regarding how LLMs work, as well as the steps of training an LLM.
The last and third lecture of the day focused on the preprocessing pipelines for AI-ready data. We identified the type of data that is generally used for training and what is the most time-consuming part of training an AI. This set of lectures also ended with a hands-on session.
Day 3, Wednesday 23 September
The third day of the summer school was fully focused on LLMs. The first lecture was about adapting and fine-tuning LLMs for low-resource languages with Leon Voukoutis. We learned about what low-resource languages are and how the training pipeline is set up for these languages.
The second lecture was about the evaluation of LLMs in downstream tasks with Prokopis Prokopidis, and the third and last lecture focused on LLMs for translation or Machine Translation in the age of LLMs by Sokratis Sofianopoulos. The day also ended with a very useful hands-on session in Google Colab, where we were exposed to more practical ways of Machine Translation.
Day 4, Thursday 24 September
The fourth day was also full of learnings concerning AI, but this time the focus was on the application of AI towards specific areas. The first lecture was about AI Application for Language Analysis by Stergios Chatzikyriakidis, while the second lecture focused on AI and low-resource language varieties, lectured by Stella Markantonatou / Stavros Bompolas / Vivian Stamou / Antonis Dimakis, while the third was AI for the study of Ancient Greek by Paraskevi Platanou, and the fourth focused on AI for Historical Document Analysis: from Images to Knowledge by Basilis Gatos. In short, all these lectures explained a more practical use of AI tools applied to these specific areas.
Day 5, Friday 25 September
The last day of the summer school had the first lecture focused on social biases in the era of LLMs by Katerina Gkirtzou. Considering my own research on language biases and attitudes, this topic was particularly interesting for me. As we all know, LLMs are prone to biases, since they are trained on human data, which also contains biases and flaws. While it might not be fully possible to eradicate these biases at the current stage, the question that remains is what are the techniques for mitigating these biases.
The second lecture was about Introduction to AI and Mechanistic Interpretability by George Paraskevopoulos. In this lecture, we were briefly led through the introduction to AI, since this is a topic that we have dived deeper into at the beginning of the summer school, and then mostly focused on Mechanistic Interpretability. We explored concepts such as explainability vs. interpretability, and the mechanistic of interpretability. This session also ended with a Jupyter Notebook in Google Colab.
The very last lecture of this summer school was presented by Alexandros Nousias and explored data streams as assets of trustworthy AI. At the beginning, we brainstormed with many questions related to the definition of trustworthy AI, and the lecture was mainly focused on Data and Input.
This summer school allowed me to deepen my knowledge in the fields of AI and ML, as well as their practical applications that can be extended beyond my field of expertise. I am definitely recommending this summer school to everyone who wants to strengthen their knowledge in AI, specifically concerning LLMs and ML. What is good about this summer school is that it is suitable for different levels of expertise, from beginner to advanced. In case you are a complete beginner, do not worry: the lecturers (and your colleagues too) will help you out with the basic setup.
We were also advised that the next edition of the summer school, previewed for 2027, will be the last. So, in case you want to apply, you still have time until November 30 to register. You can find more information on the ATRIUM TNA website.
Image 3. Last picture with the whole group and the organizers. Photo credit: Iro Tsiouli