2607.10020

Total: 1

#1 Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora [PDF] [Copy] [Kimi2] [REL]

Authors: Lars Henry Berge Olsen, Pierre Lison, Martin Jullum, Mark Anderson

We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus. The tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprints. By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities. This makes FindMyText particularly suited for verifying the presence of copyrighted material in a corpus. Leveraging a distributed, disk-based indexing framework, the system scales to large web-crawled datasets. Using a new benchmark for evaluating text containment methods, we show that FindMyText outperforms alternative approaches across three datasets (ArXiv papers, Wikipedia, and generic web content).

Subjects: Computation and Language , Artificial Intelligence

Publish: 2026-07-10 22:46:11 UTC