Categories
Meetings

Context Is Not a Feature:What AI Systems Fail to Learn

S. Mahed Mousavi from University of Trento will present his research from his PhD (first part) and more recent projects (second part) including the CORE research program.

Abstract
Conversational AI systems can now produce fluent, human-like text and are increasingly used in everyday and sensitive settings. This talk argues that fluency should not be mistaken for contextual understanding, and that current AI systems are not trained to track who they are talking to, what has happened before, or what is socially appropriate in a given situation. 

The talk presents two lines of work addressing this problem. The first focuses on deploying ConvAI models for longitudinal dialogues, i.e. a sparse sequence of personal dialogue sessions. It presents methods for collecting multi-session conversations, building and updating a model of the individual user over time, and evaluating AI models in a standardized way. These efforts resulted in the first registered Randomized Controlled Trial (RCT) with a ConvAI system in the mental health domain. The second line of work asks why AI systems struggle to understand context in the first place. It shows that many benchmarks used to claim reasoning abilities in AI are themselves flawed, and traces the problem to how these systems are trained, showing evidence that the current training paradigm is not sustainable, and that loss optimization is not a reliable signal of learning.

Where: Sala Conferenze, 3rd Floor
When: 22/09/2026, 14:00

Links:

Categories
Meetings

The role of priming and predictability in human and language model production choice

Arabella Sinclair, Assistant Professor at UCL (London) and at the University of Aberdeen will present her work who will present her work on large language models (LLMs), language processing, and cognition.

Abstract
Both humans and Large Language Models (LMs) generate predictions about upcoming words and structures based on recent context. In dialogue, speakers continuously choose how to realise their communicative intentions, with these choices shaped by multiple, sometimes competing pressures, including production costs borne by the speaker and comprehension costs incurred by the listener. Production costs reflect the effort involved in planning and generating an utterance, while comprehension costs reflect how easily an utterance can be processed and interpreted. These costs are influenced by a range of factors, including priming, utterance length, informational content, and contextual predictability.
In this talk, I explore how LMs can serve both as tools for studying human interaction and as components of broader cognitive models of language processing. First, I present evidence for parallels between humans and LMs in primed comprehension facilitation, showing how prior exposure to linguistic structures influences subsequent processing. Second, I discuss ongoing work investigating analogous effects in a controlled production setting. Finally, I introduce a new procedure for constructing contextual alternative sets that enables probabilistic pragmatic models of language production to be instantiated and evaluated at scale in open-ended communicative settings. This framework also provides a principled basis for comparing competing notions of communicative cost.
Taken together, these findings shed light on the mechanisms underlying LMs’ in-context learning behaviour while also assessing their potential as models of human linguistic processing and communicative decision-making.

When: 11/09/2026, 10:30
Where: Sala Conferenze, 3rd Floor

Categories
Meetings

Large Language Models in Emergency Medicine: Findings and Future Prospects

Bernardo Magnini, Senior Researcher at Fondazione Bruno Kessler (FBK) in Trento, will present a series of results from the eCREAM (enabling Clinical Research in Emergency and Acute care Medicine) Horizon project.

Abstract
This seminar provides an overview of the use of Large Language Models (LLMs) within the eCREAM (enabling Clinical Research in Emergency and Acute care Medicine) Horizon project. I will first introduce a unique corpus of clinical notes collected from the emergency departments of several Italian hospitals and show how these data have been used to train small LLMs for a range of clinical tasks. I will then focus on Case Report Form (CRF) filling, a key task addressed in eCREAM, where manually annotated clinical notes are leveraged to extract clinically relevant information. I will present results showing that substantial performance improvements can be achieved by fine-tuning models with additional training data automatically extracted from unannotated clinical notes. Finally, I will discuss ongoing work on the automatic interpretation of patient–doctor interactions in emergency departments, highlighting current challenges and outlining future research directions.

Where: Sala Conferenze, 3rd Floor
When: 11/09/2026, 9:30

Categories
Meetings

Context Is Not a Feature: What AI Systems Fail to Learn

Seyed Mahed Mousavi from University of Trento will give a seminar to our Department.

Abstract
Conversational AI systems can now produce fluent, human-like text and are increasingly used in everyday and sensitive settings. This talk argues that fluency should not be mistaken for contextual understanding, and that current AI systems are not trained to track who they are talking to, what has happened before, or what is socially appropriate in a given situation. 

The talk presents two lines of work addressing this problem. The first focuses on deploying ConvAI models for longitudinal dialogues, i.e. a sparse sequence of personal dialogue sessions. It presents methods for collecting multi-session conversations, building and updating a model of the individual user over time, and evaluating AI models in a standardized way. These efforts resulted in the first registered Randomized Controlled Trial (RCT) with a ConvAI system in the mental health domain. The second line of work asks why AI systems struggle to understand context in the first place. It shows that many benchmarks used to claim reasoning abilities in AI are themselves flawed, and traces the problem to how these systems are trained, showing evidence that the current training paradigm is not sustainable, and that loss optimization is not a reliable signal of learning.

When: Tuesday, 22nd of September, 14:00
Where: Sala Conferenze, 3rd floor

Categories
Meetings

The Ethics of Automated Counterspeech

Gavin Abercrombie will give a talk while visiting our group on his recent research.

Abstract
In this talk, I investigate the question of whether we should, ethically, automate counterspeech. I discuss criteria for the ethical assessment of automating counterspeech, look at the legal and philosophical justifications, both for and against doing so, and discuss how the applicability of these ethical considerations depends on automated counterspeech’s autonomy, transparency, and intended aims. I conclude by drawing together implications for both practice and future research.

Bio: Gavin Abercrombie is Assistant Professor at Heriot-Watt University, Scotland. He was a co-organiser of the Workshop on CounterSpeech for Online Harms (2023) and the ACL Tutorial on Counterspeech against Hate and Misinformation, and is a co-author of a chapter on counterspeech in the Bloomsbury Handbook of Dangerous Speech (Forthcoming).

When: Wednesday, 9th of September, 11:00
Where: Sala Conferenze, 3rd floor

Slides PDF

Categories
Meetings

What’s going on with Statistics in NLP

Samuele D’Avenia will provide a brief overview on statistical methods commonly used in NLP, with a focus on common errors.

Abstract
Statistical significance testing is widely used in Natural Language Processing, yet its application is often incomplete or misleading.
This talk provides a high-level overview of common statistical tools employed in NLP research and highlights several recurring issues that can compromise the interpretation of experimental results.

We will mainly focus on hypothesis testing, highlighting the multiple comparison and the large sample issues, while also discussing current practices.

Additionally, we will provide a brief overview of the current practices within the Italian NLP community.

When: Friday, 10th of July, 11:30
Where: Sala Conferenze, 3rd floor

Categories
Meetings

Advocacy Radar: Monitoring Advocacy for Digital Rights

Giulio Corradi (Privacy Network), and Marta Marchiori Manerba (University of Turin) will present one of their latest work.

Abstract:
In this talk we present Advocacy Radar, a platform that allows civil society practitioners, journalists, and policy researchers to monitor the state of digital rights advocacy in Italy and Europe from an interactive interface.
Advocacy Radar adapts NLP methods to collect and process heterogeneous public sources – including regulatory authorities, NGOs, and technology media – spanning both Italian and multilingual documents.
Analyses are presented through a dynamic multi-view dashboard. A human-in-the-loop correction through the interface allows domain experts to refine model outputs and incrementally improve models over time.
In this seminar, we describe the system architecture, its main dashboard components, and representative use cases for different user profiles.

Remark: The project is developed in partnership with Privacy Network, an association committed to safeguarding digital rights in Italy and beyond.
The work is ongoing; a version of this paper has been submitted to CLiC-it 2026.
We welcome informal feedback and suggestions from the CCC audience: your expertise and perspective would be super precious in shaping the next stages of development.

Short Bio:
Giulio Corradi is a Management Engineer based in Milan, working as a consultant at SDG Group. He holds an MSc in Management Engineering from the University of Bergamo, with a thesis on algorithmic accountability under the EU AI Act, developed in collaboration with the Department of Law.

His interests sit at the intersection of data-driven analysis, NLP, and business processes, with a particular focus on the social and political implications of technology. Currently He is a Research Officer at Privacy Network, one of Italy’s leading digital rights organizations, he researches automated surveillance systems and their impact on fundamental rights and democratic processes. He also writes for Scomodo, contributing to public discourse on technology, power, and society.

Marta Marchiori Manerba is a Postdoc at the Computer Science Department of the University of Turin, where she works on perspectivist approaches to dialogue modeling. She holds a Ph.D. in AI & Society from the University of Pisa, with a dissertation focused on fairness auditing through explainability particularly in the context of abusive language detection.

When: Friday, 19th of June, 11:30
Where: Sala Riunioni, 1st floor

Categories
Meetings

From Crowdsourcing to Large Language Models: Advancing Turkish Word Sense Disambiguation

The seminar is presented by Dr Dilara Torunoğlu Selamet, Lecturer in the Department of Computer Engineering at Istanbul Technical University (ITU), Türkiye.

Abstract:
Word Sense Disambiguation (WSD) is a fundamental Natural Language Processing (NLP) task that aims to identify the intended meaning of an ambiguous word in context. While substantial progress has been achieved for high-resource languages, WSD remains challenging for morphologically rich and low-resource languages such as Turkish due to the scarcity of large-scale annotated datasets.
In this talk, I will present my PhD research on Turkish Word Sense Disambiguation, which combines data-centric and model-centric approaches. First, I will introduce DodoMe, a large-scale gamified crowdsourcing platform developed to collect sense-annotated Turkish sentences. The resulting dataset contains more than 158,000 annotations covering 30 highly ambiguous Turkish words and represents one of the largest publicly available WSD resources for Turkish.
Second, I will discuss a systematic comparison of modern approaches to WSD, including contextual embedding-based methods, prompting-based inference with Large Language Models (LLMs), and instruction-based fine-tuning of open-source LLMs. The results demonstrate that recent LLMs substantially outperform traditional embedding-based approaches and that instruction tuning can further improve performance when sufficient high-quality annotated data is available.
Finally, I will discuss the broader implications of combining human computation, crowdsourcing, and large language models for developing semantic resources and language technologies for under-resourced languages.

Short Bio:
Dilara Torunoğlu Selamet is a Lecturer in the Department of Computer Engineering at Istanbul Technical University (ITU), Türkiye, and recently defended her PhD in Computer Engineering. She is a member of the ITU Natural Language Processing Research Group, and her research focuses on Natural Language Processing, lexical semantics, word sense disambiguation, meaning representation, and multilingual language technologies.
Her doctoral research investigates Turkish Word Sense Disambiguation through the integration of large-scale crowdsourced datasets and Large Language Models. She is the creator of DodoMe, a gamified crowdsourcing platform designed for collecting semantic annotations in Turkish. Her work explores contextual embeddings, prompting strategies, and instruction-based fine-tuning approaches for semantic disambiguation tasks.
Dilara is actively involved in international collaborations through the UniDive COST Action, contributing to multilingual and multimodal language technology initiatives. She serves as a language leader and coordinator in the AdMiRe (Advancing Multimodal Idiomaticity Representation) shared task series, which focuses on multilingual and multimodal idiomaticity understanding across dozens of languages.
Her recent work includes contributions to the AdMiRe shared tasks at EACL 2026 and LREC-COLING 2026, as well as research on Turkish Word Sense Disambiguation, Abstract Meaning Representation (AMR), Uniform Meaning Representation (UMR), and multilingual idiomaticity understanding. She has co-authored large-scale international publications involving researchers from more than 30 languages and actively contributes to the development of multilingual benchmarks, linguistic resources, and evaluation campaigns for Natural Language Processing.

When: 10/06/2026, h 11:00 am
Where: Sala Conferenze, third floor

Categories
Meetings

The problem of Long-tail entities: definition and benchmarking

Lia will be presenting one of her latest work.

Abstract:
Automatic verbalization of structured knowledge is a key task for making knowledge graphs accessible to non-expert users and for supporting retrieval-augmented generation systems. Although recent advances in Data-to-Text generation have improved multilingual coverage, little attention has been paid to potential biases in the verbalization of rare entities, often referred to as long-tail entities. In this work, we present the first systematic study of long-tail entities in Data-to-Text generation. We introduce TailNLG, a new multilingual benchmark in English, Italian, and Spanish, built from Wikidata and covering entities with varying levels of popularity. We evaluate three different families of large language models in zero-shot settings and compare their performance on rare versus common entities, as well as against the established WebNLG benchmark.

Since a univocal definition of long tail remains an open challenge, I will also present the first step of a research line aimed at clearly identifying the sociodemographic variables that tend to characterize underrepresented entities. This also aims to understand whether LLMs produce different outputs when dealing with well-known versus less common entities, and what implications this may have in everyday life.

Categories
Meetings

Pushing the Boundaries of Fixed Budget Best Arm Identification

CCC Seminar by Luigi Sauro, Professor of Computer Science at Università degli Studi di Napoli “Federico II”.

Abstract (Italian): Supponete di avere un insieme finito di meccanismi stocastici (arms), ognuno dei quali restituisce ad ogni interazione un reward sulla base di una distribuzione a voi sconosciuta. Lo scopo del Fixed-Budget Best Arm Identification è quello di identificare, sulla base di un numero prefissato di possibili interazioni sequenziali, l’arm con expected reward massimo. Questo problema induce un tipico dilemma exploration vs exploitation che si riscontra in molti contesti applicativi (clinical trials, wireless network selection, recommender systems, A/B testing). In questo seminario illustrerò una nuova strategia di interazione che si applica sotto le ipotesi che gli arm siano governati da distribuzioni sub-gaussiane e che fa uso di tecniche di non-linear convex optimization. Mostrerò inoltre un upper bound teorico che evidenzia il perché questa strategia raffina lo stato dell’arte ed un’analisi sperimentale che supporta la sua efficacia.

Bio: https://www.docenti.unina.it/teacher/4c55494749534155524f5352414c475537345432394638333941/profile/references

When: June 5th, 15:00
Where: Sala Riunioni, Primo Piano