Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision–Language Models

Published in Preprint — under review, 2026

Across 25 large vision–language models, we investigate unusually large activations in visual tokens. We identify a trigger direction and a rule for locating tokens likely to spike, and show that common corruptions and small adversarial perturbations can create, relocate, or remove these spikes. A preventive intervention suppresses spikes by removing the trigger component before they form, while leaving other image tokens nearly unchanged.

Read the paper on arXiv.

Recommended citation: Ngnawé, J., Pequignot, Y., Sahoo, S., Gagné, C., Precioso, F., & Koyejo, S. (2026). Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision–Language Models. arXiv preprint arXiv:2609.32808.
Download Paper