Interpreting Text-Enriched Dual-Head Multitask Learning for Indonesian Hateful Meme Detection Using Explainable AI

Authors

  • Selamet Riadi Universitas Teknologi Mataram
  • Emi Suryadi Universitas Teknologi Mataram
  • Muhamad Masjun Efendi Universitas Teknologi Mataram
  • Bahtiar Imran Universitas Teknologi Mataram
  • Muhammad Zamroni Uska Universitas Hamzanwadi

DOI:

https://doi.org/10.69916/jkbti.v5i3.576

Keywords:

Hateful Memes, Hate Speech Detection, Multitask Learning, Explainable AI, IndoBERTweet

Abstract

Internet memes in Indonesia are frequently weaponized to disseminate implicit
hate speech through sarcasm and cultural nuances. Automatically detecting
such content is computationally challenging, and existing deep learning
frameworks predominantly operate as opaque black boxes, lacking decision
transparency. This study implements and optimizes a text-enriched dual-head
multitask learning architecture utilizing IndoBERTweet to concurrently classify
hatefulness and appropriateness within the INDOMEME dataset. Rather
than processing raw image pixels, we employ a text-enrichment strategy where
visual semantics are transcribed into textual descriptors via Optical Character
Recognition and vision-language captioning. To bridge the interpretability gap,
we deploy Local Interpretable Model-agnostic Explanations (LIME) to decode
the internal feature attributions of the architecture. Furthermore, advanced
training optimizations, encompassing cosine annealing, gradient accumulation,
class-weighted loss, and dynamic threshold calibration, were engineered to
enhance model generalization. Experimental evaluations demonstrate that
the optimized model achieves a Macro-F1 score of 0.812 for hatefulness and
0.820 for appropriateness, surpassing the established baseline. Crucially, the
LIME analysis unveils a pivotal finding: despite sharing an identical textual
backbone, the hate-specific head predominantly focuses on lexicons carrying
social agitation, whereas the appropriateness head prioritizes general norm
violations. These empirical findings substantiate that multitask learning enriches
semantic representation quality, offering a transparent framework for
trustworthy content moderation.

Downloads

Download data is not yet available.

References

E. W. Pamungkas, C. S. Wahyuni, I. Amal, D. Purworini, and B. S. Rintyarna, “Decoding hate in memes: multimodal and multitask approaches for low-resource Indonesian social media,” PeerJ Comput. Sci., vol. 12, p. e3736, 2026.

M. O. Ibrohim and I. Budi, “Hate speech and abusive language detection in Indonesian social media: Progress and challenges,” Heliyon, vol. 9, no. 8, 2023.

F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A pretrained language model for Indonesian Twitter with effective domain-specific vocabulary initialization,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 10660–10668.

H. Zhang and M. O. Shafiq, “Survey of transformers and towards ensemble learning using transformers for natural language processing,” J. big Data, vol. 11, no. 1, p. 25, 2024.

J. Zhang, K. Yan, and Y. Mo, “Multi-task learning for sentiment analysis with hard-sharing and task recognition mechanisms,” Information, vol. 12, no. 5, p. 207, 2021.

M. S. Hee, W.-H. Chong, and R. K.-W. Lee, “Decoding the underlying meaning of multimodal hateful memes,” arXiv Prepr. arXiv2305.17678, 2023.

S. Ali et al., “Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence,” Inf. fusion, vol. 99, p. 101805, 2023.

A. Nirmal, A. Bhattacharjee, P. Sheth, and H. Liu, “Towards interpretable hate speech detection using large language model-extracted rationales,” in Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024), 2024, pp. 223–233.

H. Kibriya, A. Siddiqa, W. Z. Khan, and M. K. Khan, “Towards safer online communities: Deep learning and explainable AI for hate speech detection and classification,” Comput. Electr. Eng., vol. 116, p. 109153, 2024.

D. Mittal and H. Singh, “Enhancing hate speech detection through explainable ai,” in 2023 3rd International conference on smart data intelligence (ICSMDI), IEEE, 2023, pp. 118–123.

G. Burbi, A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo, “Mapping memes to words for multimodal hateful meme classification,” in 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), IEEE, 2023, pp. 2824–2828.

E. Fetahi, A. Susuri, M. Hamiti, Z. Kastrati, E. Canhasi, and A. Misini, “Enhancing social media hate speech detection in low-resource languages using transformers and explainable AI,” Soc. Netw. Anal. Min., vol. 15, no. 1, p. 82, 2025.

L. Cao, X. Liu, and H. Shen, “Adaptable focal loss for imbalanced text classification,” in International Conference on Parallel and Distributed Computing: Applications and Technologies, Springer, 2021, pp. 466–475.

O. V. Johnson, C. Xinying, K. W. Khaw, and M. H. Lee, “ps-CALR: periodic-shift cosine annealing learning rate for deep neural networks,” IEEE access, vol. 11, pp. 139171–139186, 2023.

T. Dehdarirad, “Evaluating explainability in language classification models: A unified framework incorporating feature attribution methods and key factors affecting faithfulness,” Data Inf. Manag., p. 100101, 2025.

P. Kapil and A. Ekbal, “A transformer based multi task learning approach to multimodal hate speech detection,” Nat. Lang. Process. J., vol. 11, p. 100133, 2025.

E. Hwang and V. Shwartz, “Memecap: A dataset for captioning and interpreting memes,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1433–1445.

H. Ren, Y. Zhao, Y. Zhang, and W. Sun, “Learning label smoothing for text classification,” PeerJ Comput. Sci., vol. 10, p. e2005, 2024.

C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning (still) requires rethinking generalization,” Commun. ACM, vol. 64, no. 3, pp. 107–115, 2021.

A. Rawat, S. Kumar, and S. S. Samant, “Hate speech detection in social media: Techniques, recent trends, and future challenges,” Wiley Interdiscip. Rev. Comput. Stat., vol. 16, no. 2, p. e1648, 2024.

Downloads

Published

2026-09-01

PlumX Metrics

Scite Metrics

Altmetric

How to Cite

[1]
S. Riadi, Emi Suryadi, Muhamad Masjun Efendi, Bahtiar Imran, and Muhammad Zamroni Uska, “Interpreting Text-Enriched Dual-Head Multitask Learning for Indonesian Hateful Meme Detection Using Explainable AI”, JKBTI, vol. 5, no. 3, pp. 585–593, Sep. 2026.