Domain Pre-training Impact on Representations

2025.findings-emnlp.1201@ACL

Total: 1

#1 Domain Pre-training Impact on Representations [PDF] [Copy] [Kimi] [REL]

Authors: Cesar Gonzalez-Gutierrez, Ariadna Quattoni

This empirical study analyzes how the choice of pre-training corpus affects the quality of learned transformer representations. We focus specifically on the representation quality achieved through pre-training alone. Our experiments demonstrate that pre-training on a small, specialized corpus can produce effective representations, and that the effectiveness of combining a generic and a specialized corpora depends on the distributional similarity between the target task and the specialized corpus.

Subject: EMNLP.2025 - Findings