Intricacies of Feature Geometry in Large Language Models

Ut3ml7Hdwx@OpenReview

Total: 1

#1 Intricacies of Feature Geometry in Large Language Models [PDF¹] [Copy] [Kimi²] [REL]

Authors: Satvik Golechha, Lucius Bushnaq, Euan Ong, Neeraj Kayal, Nandi Schoots

Studying the geometry of a language model's embedding space is an important and challenging task because of the various ways concepts can be represented, extracted, and used. Specifically, we want a framework that unifies both measurement (of how well a latent explains a feature/concept) and causal intervention (how well it can be used to control/steer the model). We discuss several challenges with using some recent approaches to study the geometry of categorical and hierarchical concepts in large language models (LLMs) and both theoretically and empirically justify our main takeaway, which is that their orthogonality and polytopes results are trivially true in high-dimensional spaces, and can be observed even in settings where they should not occur.

Subject: ICLR.2025 - Poster

Ut3ml7Hdwx@OpenReview

#1 Intricacies of Feature Geometry in Large Language Models [PDF1] [Copy] [Kimi2] [REL]

#1 Intricacies of Feature Geometry in Large Language Models [PDF¹] [Copy] [Kimi²] [REL]