Skip to content

MMDiff a new framework for isolating and controlling multimodal artificial intelligence features

Share
MMDiff a new framework for isolating and controlling multimodal artificial intelligence features

Listen to this article

Read by Anchor

A research paper submitted via the arXiv platform on 10 August 2026 and accepted for publication in the workshop “Trustworthy AI for Good” at ICML 2026 introduced a new framework called “MMDiff” that aims to discover and control internal features within multimodal large language models. The research team, consisting of Honar Patra, Latchin Nagashyar, Ashkan Khakzar, Philippe Tour, Christian Schroder de Vitte, Konstantin Vinhof, and Ronald Clark, explained that multimodal models possess strong visual understanding capabilities, but the internal features that give rise to these behaviors remain difficult to pinpoint, examine, or control. Although decomposing hidden states into interpretable feature directions using rare autoencoders is suitable for probing models after training, it does not readily isolate features altered by multimodal training, nor does it provide a direct tool for targeted control.

The MMDiff framework relies on training multimodal rare autoencoders and converting them into feature-level APIs for discovering model behaviors and steering their outputs. The new framework offers three primary usage modes that enable researchers to work with discovered features in a causally precise manner.

The three uses are: isolating features by comparing the rare autoencoder of the base language model with its multimodal fine-tuned counterpart to identify features that changed due to visual and linguistic training; uncovering task-specific features by analyzing the differential activation of each causal feature bottleneck token; and performing feature-level control by either removing discovered feature directions or causally steering them. To evaluate the framework’s success, the authors trained multimodal rare autoencoders for three prominent model families, LLaVA-MORE, PaliGemma 2, and InternVL 3.5, and assessed performance on three defined domains: spatial visual understanding, multimodal safety, and optical character recognition.

The experimental evaluation revealed the discovery of rare, causally specific features, and selectively removing these features reduced targeted behaviors by an average of 12 % on spatial tasks and 17 % on OCR tasks. The same feature removal also lowered the success rate of attacks on multimodal model safety by 24 %, without any negative impact on model performance in visual question-answering. Conversely, steering the identified feature directions improved spatial understanding accuracy by an average of 3.6 % and OCR accuracy by an average of 1.8 % compared with standard single-layer steering methods.

These findings, classified within computer vision, artificial intelligence, natural language processing, and machine learning, confirm that multimodal rare autoencoders are not merely tools for interpreting model behavior and dissecting architecture, but constitute genuine mechanisms for operating, auditing, and monitoring multimodal large language models to ensure their safety and to make their outputs more secure and reliable.

Don't miss the next story

Subscribe for updates