ICADL 2007 - LNCS 4822

Organizing News Archives by Near-Duplicate Copy Detection in Digital Libraries

Hung-Chi Chang¹ and Jenq-Haur Wang²

¹Institute of Information Science, Academia Sinica, Taiwan
hungchi@iis.sinica.edu.tw

²Department of Computer Science and Information Engineering, National Taipei University of Technology, Taiwan
jhwang@csie.ntut.edu.tw

Abstract. There are huge numbers of documents in digital libraries. How to effectively organize these documents so that humans can easily browse or reference is a challenging task. Existing classification methods and chronological or geographical ordering only provide partial views of the news articles. The relationships among news articles might not be easily grasped. In this paper, we propose a near-duplicate copy detection approach to organizing news archives in digital libraries. Conventional copy detection methods use word-level features which could be time-consuming and not robust to term substitutions. In this paper, we propose a sentence-level statistics-based approach to detect near-duplicate documents, which is language independent, simple but effective. It’s orthogonal to and can be used to complement word-based approaches. Also it’s insensitive to actual page layout of articles. The experimental results showed the high efficiency and good accuracy of the proposed approach in detecting near-duplicates in news archives.

Keywords: Near-duplicate document copy detection, sentence-level features, news archive organization

LNCS 4822, p. 410 ff.

Full article in PDF | BibTeX