KDD 2026 · Dataset & Benchmark

MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation

Memes cross borders. Meaning does not always follow.

MemeBridge is a culturally grounded dataset and benchmark for studying how U.S.-originated memes are interpreted—and misinterpreted—across U.S. and Chinese perspectives.

Hangxiao Zhu1 Suliu Qin2 Zhuoyan Li3 Ming Jiang4 Yu Zhang1 Meng Xia1

1Texas A&M University 2Independent Researcher 3Purdue University 4University of Wisconsin–Madison

Validated Memes
621
Cross-Cultural Reviews
1,242
Cultural-Knowledge Memes
529
Multimodal LLMs
4

The cultural gap runs both ways

Memes compress humor, history, slang, and social norms into a single image. Without the right cultural context, what feels obvious to one audience can become confusing—or mean something entirely different—to another.

MemeBridge brings two complementary views together: how Chinese participants interpret U.S.-originated memes, and how accurately U.S. contributors anticipate those cross-cultural misunderstandings. This framing reveals a two-sided gap in both interpretation and perspective-taking.

Bidirectional framing

Pairs audience interpretation with a native contributor’s anticipated misunderstanding.

Rich supervision

Adds explanations, misunderstandings, sentiment, emotion, topic, and knowledge type.

Model probing

Tests multimodal LLMs under default, U.S., and Chinese cultural role prompts.

Three stages, one culturally grounded benchmark

Three-stage MemeBridge pipeline covering collection and cleaning, U.S. validation, and cross-cultural labeling.
MemeBridge moves from 1,000 U.S. meme submissions through quality filtering and validation to 621 memes and 1,242 cross-cultural reviews.

Stage 1

Collect & clean

100 U.S. contributors submitted 1,000 memes. BERT-based filtering, lexical-diversity checks, text standardization, and translation retained 754 candidates.

Stage 2

Validate

Four U.S. raters per meme assessed explanations, misunderstandings, sentiment, emotion, and cultural significance, yielding 621 final memes.

Stage 3

Test cross-culturally

63 Chinese participants provided two reviews per meme across interpretation, sentiment, and emotion tasks, producing 1,242 responses.

621 final memes 529 cultural knowledge 92 general knowledge CC BY 4.0

Misunderstanding is difficult to anticipate

U.S. contributors did not reliably anticipate how Chinese participants would misinterpret the memes.

Across all multiple-choice responses, participants selected the GPT-4-generated distractor more often than the misunderstanding anticipated by U.S. contributors.

GPT-4-generated distractor
24.0%
Contributor-anticipated distractor
17.1%

Difference in selection rates: p < .001

Pigeon meme with a native explanation, a GPT-4 interpretation, and a human-annotated potential misunderstanding.
Based meme with a native explanation, a GPT-4 interpretation, and a human-annotated potential misunderstanding.
Well that sucks meme with a native explanation, a GPT-4 interpretation, and a human-annotated potential misunderstanding.

Cultural grounding helps—unevenly

MemeBridge evaluates Qwen, GLM, LLaMA, and GPT across explanation, multiple-choice interpretation, sentiment, and emotion. Qwen, GLM, and GPT were also fine-tuned on the dataset.

Explanation similarity scores before and after fine-tuning for GPT, Qwen, and GLM.
Explanation similarity rises for GPT and GLM, while Qwen’s scores decline slightly.
Multiple-choice accuracy before and after fine-tuning for GPT, Qwen, and GLM.
Multiple-choice accuracy improves for GPT and GLM; Qwen remains strongest despite a small decline.
Emotion accuracy before and after fine-tuning for GPT, Qwen, and GLM.
Emotion accuracy improves for Qwen and GLM but drops for GPT on the held-out test set.

Fine-tuning helps most where a model is weak.

Improvements are task-dependent. Degradation on already-strong tasks is consistent with possible overfitting on a small dataset, rather than evidence of one proven cause.

The gap is two-sided

Interpretation and perspective-taking fail in different but connected ways.

Origin is not destiny

Model behavior does not divide cleanly by the country where a model was developed.

Role prompts matter

U.S. and Chinese personas change model behavior unevenly across tasks and metrics.

BibTeX

@inproceedings{zhu2026memebridge,
  title={MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation},
  author={Zhu, Hangxiao and Qin, Suliu and Li, Zhuoyan and Jiang, Ming and Zhang, Yu and Xia, Meng},
  booktitle={Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1},
  pages={2947--2958},
  year={2026}
}