Skip to content

Computational provenance experiments succeed in embedding evidence of internal model states into generated text

Share
Computational provenance experiments succeed in embedding evidence of internal model states into generated text

Listen to this article

Read by Anchor

In a new research paper, researcher Benjamin Pillet introduces an experimental framework for computational provenance, addressing an inherent limitation in language model architectures where generated text outputs provide no independently verifiable evidence of the internal states and computations that produced them. The study examines whether detectable evidence can be embedded into generated text to decisively prove which causally relevant internal states occurred during generation, even when the final output is identical across different computational paths.

To explore this concept under controlled experimental conditions, the researcher designed an experiment using two distinct neural architectures: a standard feedforward neural network and a transformer model. Both architectures were trained on the same computational task with an enforced mandatory path passing through two discrete intermediate states, allowing the systems to reach the same computational result through two distinct internal trajectories. The researcher deliberately switched between these two internal paths during inference, authenticating the activated state, and used that authenticated state to determine a subtle statistical pattern in the generated text structure that a detection system could subsequently identify and recover with precision.

Full pass on dual verification tests

Evaluation results showed that both the feedforward network and the transformer model successfully passed all 128 matched test pairs across public evaluations and independently held-out evaluations, with the signal detector recovering the statistical pattern tied to the authenticated internal state without error. The experiments also demonstrated the reproducibility and stability of this causal computation, fully replicating across five independently trained feedforward models and three independently trained transformer models.

Limits of inference in answer-only models

In contrast, a comparative experiment on an answer-only transformer revealed the model's limitations in self-reporting its trajectory, as linear probes failed to recover naturally acquired intermediate states from within the architecture. These findings provide a solid proof of concept that authenticated information about causally relevant internal states can be embedded and preserved in generated text even when the final output remains identical, establishing an empirical foundation for tracing computational provenance in artificial intelligence.

Don't miss the next story

Subscribe for updates