Varsha Suresh

Varsha Suresh

Postdoctoral Researcher

Saarland University

About

I am a postdoctoral researcher at Multimodal Language Processing Group at Max Plank Institute of Informatics headed by Prof. Dr. Vera Demberg. I am also affiliated to Department of Language Science and Technology at Saarland University. My research focuses on multimodal language modelling and human-centered machine learning.

I am particularly interested in how verbal and non-verbal signals interact, how multimodal systems reason about people and behavior, and how to build language and vision-language models that are both more grounded and more socially aware.

News

Publications
2026

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman, M Hamza Mughal, Christian Theobalt, Ashwin Ram, Jürgen Steimle, Vera Demberg

arXiv preprint, 2026. arXiv

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

Anisha Saha, Varsha Suresh, Teodora Kamova, Sophia Wiedmann, Timothy Hospedales, Vera Demberg

arXiv preprint, 2026. arXiv

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

Toshiki Nakai, Varsha Suresh, Vera Demberg

arXiv preprint, 2026. arXiv

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

Tsan Tsai Chan, Varsha Suresh, Anisha Saha, Michael Hahn, Vera Demberg

arXiv preprint, 2026. arXiv

Modeling Turn-Taking with Semantically Informed Gestures

Varsha Suresh, M Hamza Mughal, Christian Theobalt, Vera Demberg

Findings of the Association for Computational Linguistics: EACL, 2026.

2025

Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues

Varsha Suresh, M Hamza Mughal, Christian Theobalt, Vera Demberg

ACL, 2025. Oral presentation (Top 8%).

GestureCoach: Rehearsing for Engaging Talks with LLM-Driven Gesture Recommendations

Ashwin Ram, Varsha Suresh, Artin Saberpour Abadian, Vera Demberg, Jürgen Steimle

UIST, 2025.

Hybrid Multi-view Approach Towards Augmenting Large Language Models for Human Activity Recognition

Suman Bhoi, Varsha Suresh, Hsu Wynne, Lee Mong Li

ECAI, 2025.

Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition

Frances Yung, Varsha Suresh, Zaynab Reza, Mansoor Ahmad, Vera Demberg

SIGDial, 2025.

MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection

Anisha Saha, Varsha Suresh, Timothy Hospedales, Vera Demberg

arXiv preprint, 2025. arXiv

2024

An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks

Varsha Suresh, Salah Aït-Mokhtar, Caroline Brun, Ioan Calapodescu

ICASSP, 2024.

2023

Critically examining the Domain Generalizability of Facial Expression Recognition models

Varsha Suresh, Gerard Yeo, Desmond C. Ong

arXiv preprint, 2023. arXiv

2022

Shape-Based Conditional Neural Field for Wrist-Worn Change-Point Detection

Yuang Shi, Varsha Suresh, Ooi Wei Tsang

PerCom Workshops, 2022.

Using Positive Matching Contrastive Loss with Facial Action Units to mitigate bias in Facial Expression Recognition

Varsha Suresh, Desmond C. Ong

ACII, 2022.

Context-Dependent Deep Learning for Affective Computing

Varsha Suresh

ACII Workshops and Demos, 2022.

2021

Using knowledge-embedded attention to augment pre-trained language models for fine-grained emotion recognition

Varsha Suresh, Desmond C. Ong

arXiv preprint, 2021. arXiv

A systematic evaluation of domain adaptation in facial expression recognition

Yan San Kong, Varsha Suresh, Jonathan Soh, Desmond C. Ong

arXiv preprint, 2021. arXiv

Not all negatives are equal: Label-aware contrastive loss for fine-grained text classification

Varsha Suresh, Desmond C. Ong

EMNLP, 2021.

2016

A novel technique for identifying attentional selection in a dichotic environment

Priya Shree, Piyush Swami, Varsha Suresh, Tapan Kumar Gandhi

INDICON, 2016.

Teaching & Supervision

Master Students

Currently Ongoing

  • Zhanna (Mar 2025 - present): Pragmatic Incongruity in Large Vision Language Models.
  • Diana (Mar 2025 - present): Investigating Irony Understanding in Large Vision Language Models.

Research Student Experience

  • Fan Wu: Working Toward Emotion Understanding in VLMs.

Teaching

  • Multimodal Language Understanding, Saarland University, Winter 2024 and Summer 2026.
  • Previous teaching experience at NUS (Jan 2019 - Dec 2021): IS4152 Affective Computing and CS1010E Programming Methodology for Python/C.
Talks
  • Beyond Transcript: From Understanding Non-Verbal Signals to Generating Meaningful Behaviors
    RTG Neuroexplicit group, June 25, 2026.
Contact