Skip to main content

Stealthy Backdoor Attacks on CLIP via Stylistic Textual Triggers

  • Conference paper
  • First Online:
Image and Graphics (ICIG 2025)

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 16163))

Included in the following conference series:

  • 740 Accesses

Abstract

Vision-Language Models (VLMs), such as CLIP, have been widely deployed in various cross-modal applications due to their strong alignment capability across image and text domains. However, current backdoor attacks against CLIP have predominantly focused on its image encoder, while attacks targeting the text encoder remain largely unexplored. Existing textual backdoor methods primarily rely on inserting fixed phrases or tokens, which often disrupt the fluency and semantics of the original text, making them easier to detect. To address this issue, we propose a method called Stylistic Text Encoder Attack (STEA), which leverages a large language model to generate stylistically diverse text. By applying structural modifications such as emphatic constructions and appositive phrases, our method subtly embeds backdoor triggers while preserving the naturalness and readability of the text. To mitigate potential hallucinations generated by the LLM during style transformations, we introduce a Semantic Invariance Mechanism based on structural and semantic entropy, thereby filtering out transformations that significantly deviate from the original meaning. Extensive experiments demonstrate that our approach not only preserves fluency and semantic consistency but also achieves a comparable attack success rate to insertion-based backdoor attacks.

This work is supported by the Beijing Natural Science Foundation (JQ23018) and the National Natural Science Foundation of China (No. 62276257).

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Subscribe and save

Springer+
from $39.99 /Month
  • Starting from 10 chapters or articles per month
  • Access and download chapters and articles from more than 300k books and 2,500 journals
  • Cancel anytime
View plans

Buy Now

Chapter
USD 29.95
Price excludes VAT (USA)
eBook
USD 69.99
Price excludes VAT (USA)
Softcover Book
USD 89.99
Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Similar content being viewed by others

References

  1. Chen, X., et al.: BadNL: backdoor attacks against NLP models with semantic-preserving improvements. In: Proceedings of the 37th Annual Computer Security Applications Conference, pp. 554–569 (2021)

    Google Scholar 

  2. Conde, M.V., Turgutlu, K.: Clip-art: contrastive pre-training for fine-grained art classification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3956–3960 (2021)

    Google Scholar 

  3. Dai, J., Chen, C., Li, Y.: A backdoor attack against LSTM-based text classification systems. IEEE Access 7, 138872–138878 (2019)

    Article  Google Scholar 

  4. Fang, H., Xiong, P., Xu, L., Chen, Y.: Clip2video: mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097 (2021)

  5. Gao, K., et al.: Inducing high energy-latency of large vision-language models with verbose images. arXiv preprint arXiv:2401.11170 (2024)

  6. Goel, S., et al.: Cyclip: cyclic contrastive language-image pretraining. Adv. Neural. Inf. Process. Syst. 35, 6704–6719 (2022)

    Article  Google Scholar 

  7. Jia, C., et al.: Scaling up visual and vision-language representation learning with noisy text supervision. In: International conference on machine learning, pp. 4904–4916. PMLR (2021)

    Google Scholar 

  8. Jia, J., Liu, Y., Gong, N.Z.: Badencoder: Backdoor attacks to pre-trained encoders in self-supervised learning. In: 2022 IEEE Symposium on Security and Privacy (SP), pp. 2043–2059. IEEE (2022)

    Google Scholar 

  9. Kenton, J.D.M.W.C., Toutanova, L.K.: BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of naacL-HLT, vol. 1. Minneapolis, Minnesota (2019)

    Google Scholar 

  10. Kurita, K., Michel, P., Neubig, G.: Weight poisoning attacks on pre-trained models. arXiv preprint arXiv:2004.06660 (2020)

  11. Lee, J., et al.: UniCLIP: unified framework for contrastive language-image pre-training. Adv. Neural. Inf. Process. Syst. 35, 1008–1019 (2022)

    Article  Google Scholar 

  12. Li, C., et al.: An embarrassingly simple backdoor attack on self-supervised learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4367–4378 (2023)

    Google Scholar 

  13. Li, Y., et al.: Supervision exists everywhere: a data efficient contrastive language-image pre-training paradigm. arXiv preprint arXiv:2110.05208 (2021)

  14. Lin, T.Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740–755. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10602-1_48

    Chapter  Google Scholar 

  15. Opdahl, A.L., et al.: Trustworthy journalism through AI. Data Knowl. Eng. 146, 102182 (2023)

    Article  Google Scholar 

  16. Paszke, A.: Pytorch: an imperative style, high-performance deep learning library. arXiv preprint arXiv:1912.01703 (2019)

  17. Peng, F., Yang, X., Xiao, L., Wang, Y., Xu, C.: SgVA-CLIP: semantic-guided visual adapting of vision-language models for few-shot image classification. IEEE Trans. Multimedia 26, 3469–3480 (2023)

    Article  Google Scholar 

  18. Radford, A., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning, pp. 8748–8763. PMLR (2021)

    Google Scholar 

  19. Rao, Y., et al.: DenseCLIP: language-guided dense prediction with context-aware prompting. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 18082–18091 (2022)

    Google Scholar 

  20. Sang, L., Xu, M., Qian, S., Wu, X.: Adversarial heterogeneous graph neural network for robust recommendation. IEEE Trans. Comput. Soc. Syst. 10(5), 2660–2671 (2023). https://doi.org/10.1109/TCSS.2023.3268683

    Article  Google Scholar 

  21. Shen, Y., et al.: ChatGPT and other large language models are double-edged swords (2023)

    Google Scholar 

  22. Singh, A., et al.: Flava: a foundational language and vision alignment model. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 15638–15650 (2022)

    Google Scholar 

  23. Wang, Y., Xue, D., Zhang, S., Qian, S.: BadAgent: inserting and activating backdoor attacks in LLM agents. In: Ku, L.W., Martins, A., Srikumar, V. (eds.) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 9811–9827. Association for Computational Linguistics, Bangkok, Thailand (2024). https://doi.org/10.18653/v1/2024.acl-long.530, https://aclanthology.org/2024.acl-long.530/

  24. Weiser, B., Schweber, N.: Lawyer who used ChatGPT faces penalty for made up citations. The New York Times 8 (2023)

    Google Scholar 

  25. Wu, H.H., Seetharaman, P., Kumar, K., Bello, J.P.: Wav2CLIP: learning robust audio representations from clip. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4563–4567. IEEE (2022)

    Google Scholar 

  26. Xiao, Y., Wang, W.Y.: On hallucination and predictive uncertainty in conditional language generation. arXiv preprint arXiv:2103.15025 (2021)

  27. Yang, Z., et al.: XLNet: generalized autoregressive pretraining for language understanding. Adv. Neural Inf. Process. Syst. 32 (2019)

    Google Scholar 

  28. Yang, Z., et al.: Data poisoning attacks against multimodal encoders. In: International Conference on Machine Learnin,. pp. 39299–39313. PMLR (2023)

    Google Scholar 

Download references

Author information

Authors and Affiliations

Authors

Corresponding author

Correspondence to Shengsheng Qian.

Editor information

Editors and Affiliations

Rights and permissions

Reprints and permissions

Copyright information

© 2026 The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd.

About this paper

Check for updates. Verify currency and authenticity via CrossMark

Cite this paper

Cao, K., Wang, B., Qian, S. (2026). Stealthy Backdoor Attacks on CLIP via Stylistic Textual Triggers. In: Lin, Z., et al. Image and Graphics. ICIG 2025. Lecture Notes in Computer Science, vol 16163. Springer, Singapore. https://doi.org/10.1007/978-981-95-3729-7_23

Download citation

Keywords

Publish with us

Policies and ethics