Structured queries,language modeling,and relevance modeling in cross-language information retrieval期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

按检索

Structured queries,language modeling,and relevance modeling in cross-language information retrieval

Institution:	1. LIG, University of Grenoble Alpes, Inria Grenoble - Rhône-Alpes, 655 avenue de l''Europe, 38330 Montbonnot, France;2. Department of Computer Science, Middlesex University, The Burroughs Hendon, NW4 4BT, London, United Kingdom;3. Institute of Information Science and Technologies, National Research Council, via G. Moruzzi 1, 56124, Pisa, Italy;4. Department of Computer Science, University of Pisa, Largo B. Pontecorvo 3, 56127, Pisa, Italy

Abstract:	Two probabilistic approaches to cross-lingual retrieval are in wide use today, those based on probabilistic models of relevance, as exemplified by INQUERY, and those based on language modeling. INQUERY, as a query net model, allows the easy incorporation of query operators, including a synonym operator, which has proven to be extremely useful in cross-language information retrieval (CLIR), in an approach often called structured query translation. In contrast, language models incorporate translation probabilities into a unified framework. We compare the two approaches on Arabic and Spanish data sets, using two kinds of bilingual dictionaries––one derived from a conventional dictionary, and one derived from a parallel corpus. We find that structured query processing gives slightly better results when queries are not expanded. On the other hand, when queries are expanded, language modeling gives better results, but only when using a probabilistic dictionary derived from a parallel corpus.We pursue two additional issues inherent in the comparison of structured query processing with language modeling. The first concerns query expansion, and the second is the role of translation probabilities. We compare conventional expansion techniques (pseudo-relevance feedback) with relevance modeling, a new IR approach which fits into the formal framework of language modeling. We find that relevance modeling and pseudo-relevance feedback achieve comparable levels of retrieval and that good translation probabilities confer a small but significant advantage.

Keywords:
本文献已被 ScienceDirect 等数据库收录！

设为首页 | 免责声明 | 关于勤云 | 加入收藏