2501.13100

Total: 1

#1 A Rate-Distortion Framework for Summarization [PDF] [Copy] [Kimi] [REL]

Authors: Enes Arda, Aylin Yener

This paper introduces an information-theoretic framework for text summarization. We define the summarizer rate-distortion function and show that it provides a fundamental lower bound on summarizer performance. We describe an iterative procedure, similar to Blahut-Arimoto algorithm, for computing this function. To handle real-world text datasets, we also propose a practical method that can calculate the summarizer rate-distortion function with limited data. Finally, we empirically confirm our theoretical results by comparing the summarizer rate-distortion function with the performances of different summarizers used in practice.

Subjects: Information Theory , Computation and Language , Machine Learning

Publish: 2025-01-22 18:57:14 UTC