Detecting Identical Entities in the Semantic Web Data

Authors: Michal Holub, Ondrej Proksa, Maria Bielikova
Year: 2015
Venue: Proc. of SOFSEM 2015: Theory and Practice of Computer Science, Lecture Notes in Computer Science Volume 8939, 2015, pp 519-530.
Product of the Action: Yes

Keystone Members Authors:

Large amount of entities published by various sources inevitably introduces inaccuracies, mainly duplicated information. These can even be found within a single dataset. In this paper we propose a method for automatic discovery of identity relationship between two entities (also known as instance matching) in a dataset represented as a graph (e.g. in the Linked Data Cloud). Our method can be used for cleaning existing datasets from duplicates, validating of existing identity relationships between entities within a dataset, or for connecting di erent datasets using the owl:sameAs relationship. Our method is based on the analysis of sub-graphs formed by entities, their properties and existing relationships between them. It can learn a common similarity threshold for particular dataset, so it is adaptable to its di erent properties. We evaluated our method by conducting several experiments on data from the domains of public administration and digital libraries.