I am a postdoctoral researcher at Multimodal Language Processing Group at Max Plank Institute of Informatics headed by Prof. Dr. Vera Demberg. I am also affiliated to Department of Language Science and Technology at Saarland University. My research focuses on multimodal language modelling and human-centered machine learning.
I am particularly interested in how verbal and non-verbal signals interact, how multimodal systems reason about people and behavior, and how to build language and vision-language models that are both more grounded and more socially aware.
Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
arXiv preprint, 2026. arXiv
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
arXiv preprint, 2026. arXiv
Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
arXiv preprint, 2026. arXiv
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
arXiv preprint, 2026. arXiv
Modeling Turn-Taking with Semantically Informed Gestures
Findings of the Association for Computational Linguistics: EACL, 2026.
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
ACL, 2025. Oral presentation (Top 8%).
GestureCoach: Rehearsing for Engaging Talks with LLM-Driven Gesture Recommendations
UIST, 2025.
Hybrid Multi-view Approach Towards Augmenting Large Language Models for Human Activity Recognition
ECAI, 2025.
Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition
SIGDial, 2025.
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
arXiv preprint, 2025. arXiv
An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks
ICASSP, 2024.
Critically examining the Domain Generalizability of Facial Expression Recognition models
arXiv preprint, 2023. arXiv
Shape-Based Conditional Neural Field for Wrist-Worn Change-Point Detection
PerCom Workshops, 2022.
Using Positive Matching Contrastive Loss with Facial Action Units to mitigate bias in Facial Expression Recognition
ACII, 2022.
Context-Dependent Deep Learning for Affective Computing
ACII Workshops and Demos, 2022.
Using knowledge-embedded attention to augment pre-trained language models for fine-grained emotion recognition
arXiv preprint, 2021. arXiv
A systematic evaluation of domain adaptation in facial expression recognition
arXiv preprint, 2021. arXiv
Not all negatives are equal: Label-aware contrastive loss for fine-grained text classification
EMNLP, 2021.
A novel technique for identifying attentional selection in a dichotic environment
INDICON, 2016.