Please use this identifier to cite or link to this item: https://doi.org/10.1145/1835449.1835670
Title: A Co-learning framework for learning user search intents from rule-generated training data
Authors: Yan, J.
Zheng, Z.
Jiang, L.
Li, Y.
Yan, S. 
Chen, Z.
Keywords: Classification
Search engine
User intent
Issue Date: 2010
Source: Yan, J.,Zheng, Z.,Jiang, L.,Li, Y.,Yan, S.,Chen, Z. (2010). A Co-learning framework for learning user search intents from rule-generated training data. SIGIR 2010 Proceedings - 33rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval : 895-896. ScholarBank@NUS Repository. https://doi.org/10.1145/1835449.1835670
Abstract: Learning to understand user search intents from their online behaviors is crucial for both Web search and online advertising. However, it is a challenging task to collect and label a sufficient amount of high quality training data for various user intents such as "compare products", "plan a travel", etc. Motivated by this bottleneck, we start with some user common sense, i.e. a set of rules, to generate training data for learning to predict user intents. The rule-generated training data are however hard to be used since these data are generally imperfect due to the serious data bias and possible data noises. In this paper, we introduce a Co-learning Framework (CLF) to tackle the problem of learning from biased and noisy rule-generated training data. CLF firstly generates multiple sets of possibly biased and noisy training data using different rules, and then trains the individual user search intent classifiers over different training datasets independently. The intermediate classifiers are then used to categorize the training data themselves as well as the unlabeled data. The confidently classified data by one classifier are added to other training datasets and the incorrectly classified ones are instead filtered out from the training datasets. The algorithmic performance of this iterative learning procedure is theoretically guaranteed. © 2010 ACM.
Source Title: SIGIR 2010 Proceedings - 33rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
URI: http://scholarbank.nus.edu.sg/handle/10635/68735
ISBN: 9781605588964
DOI: 10.1145/1835449.1835670
Appears in Collections:Staff Publications

Show full item record
Files in This Item:
There are no files associated with this item.

Page view(s)

40
checked on Feb 16, 2018

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.