Bidirectional framing
Pairs audience interpretation with a native contributor’s anticipated misunderstanding.
KDD 2026 · Dataset & Benchmark
Memes cross borders. Meaning does not always follow.
MemeBridge is a culturally grounded dataset and benchmark for studying how U.S.-originated memes are interpreted—and misinterpreted—across U.S. and Chinese perspectives.
1Texas A&M University 2Independent Researcher 3Purdue University 4University of Wisconsin–Madison
Abstract
Memes compress humor, history, slang, and social norms into a single image. Without the right cultural context, what feels obvious to one audience can become confusing—or mean something entirely different—to another.
MemeBridge brings two complementary views together: how Chinese participants interpret U.S.-originated memes, and how accurately U.S. contributors anticipate those cross-cultural misunderstandings. This framing reveals a two-sided gap in both interpretation and perspective-taking.
Pairs audience interpretation with a native contributor’s anticipated misunderstanding.
Adds explanations, misunderstandings, sentiment, emotion, topic, and knowledge type.
Tests multimodal LLMs under default, U.S., and Chinese cultural role prompts.
Dataset
Stage 1
100 U.S. contributors submitted 1,000 memes. BERT-based filtering, lexical-diversity checks, text standardization, and translation retained 754 candidates.
Stage 2
Four U.S. raters per meme assessed explanations, misunderstandings, sentiment, emotion, and cultural significance, yielding 621 final memes.
Stage 3
63 Chinese participants provided two reviews per meme across interpretation, sentiment, and emotion tasks, producing 1,242 responses.
Findings
Across all multiple-choice responses, participants selected the GPT-4-generated distractor more often than the misunderstanding anticipated by U.S. contributors.
Difference in selection rates: p < .001
Models
MemeBridge evaluates Qwen, GLM, LLaMA, and GPT across explanation, multiple-choice interpretation, sentiment, and emotion. Qwen, GLM, and GPT were also fine-tuned on the dataset.
What the mixed result means
Improvements are task-dependent. Degradation on already-strong tasks is consistent with possible overfitting on a small dataset, rather than evidence of one proven cause.
Interpretation and perspective-taking fail in different but connected ways.
Model behavior does not divide cleanly by the country where a model was developed.
U.S. and Chinese personas change model behavior unevenly across tasks and metrics.
Citation
@inproceedings{zhu2026memebridge,
title={MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation},
author={Zhu, Hangxiao and Qin, Suliu and Li, Zhuoyan and Jiang, Ming and Zhang, Yu and Xia, Meng},
booktitle={Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1},
pages={2947--2958},
year={2026}
}