Interpreting Text-Enriched Dual-Head Multitask Learning for Indonesian Hateful Meme Detection Using Explainable AI
DOI:
https://doi.org/10.69916/jkbti.v5i3.576Keywords:
Hateful Memes, Hate Speech Detection, Multitask Learning, Explainable AI, IndoBERTweetAbstract
Internet memes in Indonesia are frequently weaponized to disseminate implicit
hate speech through sarcasm and cultural nuances. Automatically detecting
such content is computationally challenging, and existing deep learning
frameworks predominantly operate as opaque black boxes, lacking decision
transparency. This study implements and optimizes a text-enriched dual-head
multitask learning architecture utilizing IndoBERTweet to concurrently classify
hatefulness and appropriateness within the INDOMEME dataset. Rather
than processing raw image pixels, we employ a text-enrichment strategy where
visual semantics are transcribed into textual descriptors via Optical Character
Recognition and vision-language captioning. To bridge the interpretability gap,
we deploy Local Interpretable Model-agnostic Explanations (LIME) to decode
the internal feature attributions of the architecture. Furthermore, advanced
training optimizations, encompassing cosine annealing, gradient accumulation,
class-weighted loss, and dynamic threshold calibration, were engineered to
enhance model generalization. Experimental evaluations demonstrate that
the optimized model achieves a Macro-F1 score of 0.812 for hatefulness and
0.820 for appropriateness, surpassing the established baseline. Crucially, the
LIME analysis unveils a pivotal finding: despite sharing an identical textual
backbone, the hate-specific head predominantly focuses on lexicons carrying
social agitation, whereas the appropriateness head prioritizes general norm
violations. These empirical findings substantiate that multitask learning enriches
semantic representation quality, offering a transparent framework for
trustworthy content moderation.
Downloads
References
E. W. Pamungkas, C. S. Wahyuni, I. Amal, D. Purworini, and B. S. Rintyarna, “Decoding hate in memes: multimodal and multitask approaches for low-resource Indonesian social media,” PeerJ Comput. Sci., vol. 12, p. e3736, 2026.
M. O. Ibrohim and I. Budi, “Hate speech and abusive language detection in Indonesian social media: Progress and challenges,” Heliyon, vol. 9, no. 8, 2023.
F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A pretrained language model for Indonesian Twitter with effective domain-specific vocabulary initialization,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 10660–10668.
H. Zhang and M. O. Shafiq, “Survey of transformers and towards ensemble learning using transformers for natural language processing,” J. big Data, vol. 11, no. 1, p. 25, 2024.
J. Zhang, K. Yan, and Y. Mo, “Multi-task learning for sentiment analysis with hard-sharing and task recognition mechanisms,” Information, vol. 12, no. 5, p. 207, 2021.
M. S. Hee, W.-H. Chong, and R. K.-W. Lee, “Decoding the underlying meaning of multimodal hateful memes,” arXiv Prepr. arXiv2305.17678, 2023.
S. Ali et al., “Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence,” Inf. fusion, vol. 99, p. 101805, 2023.
A. Nirmal, A. Bhattacharjee, P. Sheth, and H. Liu, “Towards interpretable hate speech detection using large language model-extracted rationales,” in Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024), 2024, pp. 223–233.
H. Kibriya, A. Siddiqa, W. Z. Khan, and M. K. Khan, “Towards safer online communities: Deep learning and explainable AI for hate speech detection and classification,” Comput. Electr. Eng., vol. 116, p. 109153, 2024.
D. Mittal and H. Singh, “Enhancing hate speech detection through explainable ai,” in 2023 3rd International conference on smart data intelligence (ICSMDI), IEEE, 2023, pp. 118–123.
G. Burbi, A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo, “Mapping memes to words for multimodal hateful meme classification,” in 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), IEEE, 2023, pp. 2824–2828.
E. Fetahi, A. Susuri, M. Hamiti, Z. Kastrati, E. Canhasi, and A. Misini, “Enhancing social media hate speech detection in low-resource languages using transformers and explainable AI,” Soc. Netw. Anal. Min., vol. 15, no. 1, p. 82, 2025.
L. Cao, X. Liu, and H. Shen, “Adaptable focal loss for imbalanced text classification,” in International Conference on Parallel and Distributed Computing: Applications and Technologies, Springer, 2021, pp. 466–475.
O. V. Johnson, C. Xinying, K. W. Khaw, and M. H. Lee, “ps-CALR: periodic-shift cosine annealing learning rate for deep neural networks,” IEEE access, vol. 11, pp. 139171–139186, 2023.
T. Dehdarirad, “Evaluating explainability in language classification models: A unified framework incorporating feature attribution methods and key factors affecting faithfulness,” Data Inf. Manag., p. 100101, 2025.
P. Kapil and A. Ekbal, “A transformer based multi task learning approach to multimodal hate speech detection,” Nat. Lang. Process. J., vol. 11, p. 100133, 2025.
E. Hwang and V. Shwartz, “Memecap: A dataset for captioning and interpreting memes,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1433–1445.
H. Ren, Y. Zhao, Y. Zhang, and W. Sun, “Learning label smoothing for text classification,” PeerJ Comput. Sci., vol. 10, p. e2005, 2024.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning (still) requires rethinking generalization,” Commun. ACM, vol. 64, no. 3, pp. 107–115, 2021.
A. Rawat, S. Kumar, and S. S. Samant, “Hate speech detection in social media: Techniques, recent trends, and future challenges,” Wiley Interdiscip. Rev. Comput. Stat., vol. 16, no. 2, p. e1648, 2024.
Downloads
Published
Scite Metrics
Altmetric
How to Cite
Issue
Section
License
Copyright (c) 2026 Selamet Riadi, Emi Suryadi, Muhamad Masjun Efendi, Bahtiar Imran, Muhammad Zamroni Uska

This work is licensed under a Creative Commons Attribution 4.0 International License.
Most read articles by the same author(s)
- Diki Hananta Firdaus, Bahtiar Imran, Lalu Darmawan Bakti, Emi Suryadi, KLASIFIKASI PENYAKIT KATARAK BERDASARKAN CITRA MENGGUNAKAN METODE CONVOLUTIONAL NEURAL NETWORK (CNN) BERBASIS WEB , Jurnal Kecerdasan Buatan dan Teknologi Informasi: Vol. 1 No. 3 (2022): Desember 2022
- Nining Putri Ningsih, Emi Suryadi, Lalu Darmawan Bakti, Bahtiar Imran, KLASIFIKASI PENYAKIT EARLY BLIGHT DAN LATE BLIGHT PADA TANAMAN TOMAT BERDASARKAN CITRA DAUN MENGGUNAKAN METODE CNN BERBASIS WEBSITE , Jurnal Kecerdasan Buatan dan Teknologi Informasi: Vol. 1 No. 3 (2022): Desember 2022
- Muhamad Masjun Efendi, Ardiyallah Akbar, Lalu Mutawalli, Development of a Web-Based Health Information System Integrated with Artificial Intelligence Using Agile Methodology at Dasan Agung Primary Health Center , Jurnal Kecerdasan Buatan dan Teknologi Informasi: Vol. 5 No. 3 (2026): September 2026 In progress.







