暂无图片
暂无图片
暂无图片
暂无图片
暂无图片
21-iCrowd An Adaptive Crowdsourcing Framework.pdf
174
16页
0次
2022-03-02
免费下载
iCrowd: An Adaptive Crowdsourcing Framework
Ju Fan
Guoliang Li
Beng Chin Ooi
Kian-lee Tan
Jianhua Feng
School of Computing, National University of Singapore, Singapore.
Department of Computer Science, TNList, Tsinghua University, Beijing, China.
{fanj, ooibc, tankl}@comp.nus.edu.sg; {liguoliang,fengjh}@tsinghua.edu.cn
ABSTRACT
Crowdsourcing is widely accepted as a means for resolving
tasks that machines are not good at. Unfortunately, Crowd-
sourcing may yield relatively low-quality results if there is no
proper quality control. Although previous studies attempt
to eliminate “bad” workers by u sing qualification tests, the
accuracies estimated from qualifications may not be accu-
rate, because workers have diverse accuracies across tasks.
Thus, the quality of the results could be further improved
by selectively assigning tasks to the workers who are well ac-
quainted with the tasks. To this end, we prop ose an adaptive
crowdsourcing framework, called iCrowd. iCrowd on-the-fly
estimates accuracies of a worker by evaluating her perfor-
mance on the completed tasks, and p redicts which tasks the
worker is well acquainted with. When a worker requests for
atask,iCrowd assigns her a task, to which the work er has
the highest estimated accuracy among all online workers.
Once a worker submits an answer to a task, iCrowd analyzes
her answer and adjusts estimation of her accuracies to im-
prove subsequent task assignments. This pap er studies the
challenges that arise in iCrowd.Therstishowtoestimate
diverse accuracies of a worker based on her completed tasks.
The second is instant task assignment. We deploy iCrowd
on Amazon Mechanical Turk, and conduct extensive exper-
iments on real datasets. Experimental results show that
iCrowd achieves higher quality than existing approaches.
Categories and Subject Descriptors
H.3.3 [Information Storage and Retriev al]: Information
Search and Retrieval—Information filtering
Keywords
Crowdsourcing; Quality control; Adaptive task assignment
1. INTRODUCTION
Crowdsourcing outsources tasks for solutions from an un-
known group of people (aka workers), which is indeed useful
to many real-world applications, e.g., image search, entity
Permission to make digital or hard copies of all or part of thisworkforpersonalor
classroom use is granted without fee provided that copies arenotmadeordistributed
for profit or commercial advantage and that copies bear this notice and the full cita-
tion on the first page. Copyrights for components of this work owned by others than
ACM must be honored. Abstracting with credit is permitted. Tocopyotherwise,orre-
publish, to post on servers or to redistribute to lists, requires prior specific permission
and/or a fee. Request permissions from permissions@acm.org.
SIGMOD’15, May 31 Jun 4, 2015, Melbourne, Victoria, Australia.
Copyright
c
2015 ACM 978-1-4503-2758-9/15/05 ...$15.00.
http://dx.doi.org/10.1145/2723372.2750550.
resolution, answering database-hard queries [24, 23, 27, 28,
12, 32]. Due to its openness, crowdsourcing yields relatively
low -quality results, or even noise, which attracts great inter-
est in devising good quality control methods [17]. Existing
methods [32, 18, 30] employ a redundancy -based strategy
which publishes a crowdsourcing task to multiple work ers
and derives the result by aggregating worker answ ers. A
na
¨
ıve aggregation approach is majority voting that choos-
es the answer that the majority of the workers yield as the
result. Recently, more sophisticated approaches have been
proposed and they can be broadly classified into two cat-
egories. The gold-injected approaches [22] leverage small
amounts of tasks with ground truth to estimate workers’
quality, while the EM -based approaches [31, 8] simultane-
ously estimate worker quality and predict aggregated results
using an Expectation-Maximization (EM) strategy. More-
over, existing approaches further improve the quality by e-
liminating “bad” w orkers. A well-known way is to use quali-
fication tests to distinguish b ad and good workers, and stop
assigning tasks to workers that cannot give good answers to
the q u alification tasks.
Although existing methods perform well in simple crowd-
sourcing tasks, such as image labeling, they may have limita-
tions on more complicated crowdsourcing tasks that require
domain knowledge. In these tasks, workers may have diverse
accuracies across tasks, as they are usually good at tasks in
domains they are familiar with but may provide low-quality
answers in unfamiliar domains. Take crowdsourced entit y
resolution [32] as an example. A worker acquainted with
Samsung stands a better chance to correctly dierentiate
the mo d els “Note4” and “S4”, while she may not be good at
tasks about iPad and cannot identify that “iPad with Reti-
na display” is colloquially referred to as “iPad 4”. Workers
with dierent backgrounds may be good at dierent topics:
abasketballfanstandsabetterchancetocorrectlyanno-
tate tables related to NBA, while a film enthusiast is more
reliable for tab les involving Hollywood films. Similar obser-
vations can be found in other tasks. W e have conducted
empirical investigation on two complicated crowdsourcing
tasks, evaluating quality of Yahoo Answers and comparing
items (e.g., which car is more fuel ecient), and report em-
pirical observations of accuracy diversity in Figure 6.
The accuracy diversity in crowdsourcing engenders many
challenges, making existing solutions inadequate for produc-
ing high-quality result. On the one hand, a worker, who
gives good answ ers to qualification tests, may not provide
promising answers to other assigned tasks. As such, the ex-
isting approaches ma y over- or under-estimate workers’ ac-
Table 1: Microtasks for verifying whether two entities are matched.
Microtask Verifying two entities Tokens
t
1
(iphone 4 WiFi 32GB, iphone four 3G black) {iphone 4 WiFi 32GB four 3G black}
t
2
(ipod touch 32GB WiFi, ipod touch headphone) {ipod touch 32GB WiFi headphone}
t
3
(ipad 3 WiFi 32GB black, new ipad cover white) {ipad 3 WiFi 32GB black new cover white}
t
4
(iphone four WiFi 16GB, iphone four 3G 16GB) {iphone four WiFi 16GB 3G}
t
5
(iphone 4 case black, iphone 4 WiFi 32GB) {iphone 4 case black WiFi 32GB}
t
6
(iphone 4 WiFi 32GB, iphone four WiFi 32GB) {iphone 4 WiFi 32GB four}
t
7
(ipod touch 32GB WiFi, ipod touch case black) {ipod touch 32GB WiFi case black}
t
8
(ipod touch headphone, ipo d nano headphone) {ipod touch nano headphone}
t
9
(ipod touch WiFi, ipod n ano headphone) {ipod touch WiFi nano headphone}
t
10
(ipad 3 WiFi 32GB black, iphone 4 cover white) {ipad 3 WiFi 32GB black iphone 4 cover white}
t
11
(ipad 4 WiFi 16GB, ipad retina display WiFi 16GB) {ipad 4 WiFi 16GB retina display}
t
12
(ipad 3 cover white, n ew ipad cover white) {ipad 3 cover white new}
curacies and thus result in unreliable aggregated results. On
the other hand, existing approaches neglect a fact that we
can adaptively assign tasks to workers who have expertise on
the tasks to further improve the quality, instead of random
task assignment without considering workers’ expertise.
To address the limitations of existing approaches, we pro-
pose an adaptive crowdsourcing framework, called iCrowd.
iCrowd on-the-fly estimates accuracies of a worker by eval-
uating her performance on the completed tasks, and infers
workers accuracies on similar tasks. When a worker re-
quests for a task, the framework assigns the worker a task,
to which the worker has the highest estimated accuracy a-
mong all online workers. Once a worker submits her answer
to a task, iCrowd analyzes her answer and adaptively adjusts
the accuracy estimation to improve any subsequent task as-
signments. In this wa y, iCrowd can eectively predict which
workers are more appropriate for a task, and adaptively as-
signs the task to these high-quality workers.
We address two main research challenges that arise in
adaptive crowdsourcing. The first one is how to estimate
the diverse accuracies of workers based on their completed
tasks. To address this challenge, we prop ose an accuracy es-
timation method by considering the “similarity” of tasks: a
worker may have comparable accuracies on tasks in similar
domains. We first construct a graph to model similarity of
tasks and evaluate worker accuracies on her completed tasks.
The second one is instant task assignment based on the es-
timated accuracies. As workers are generally impatient to
wait for too long for a task assignment, we need to eciently
assign tasks to the workers. We develop ecient algorithm-
stosupportinstanttaskassignment. Sinceexistingplat-
forms, such as Amazon Mechanical Turk (AMT) [2], have
no functionality to support assigning tasks to workers, we
develop an iCrowd system which iteratively communicates
with the platforms to receive task requests from workers,
assign tasks to them, and obtain answers from the workers.
To summarize, we make the following contributions.
(1) We formulate the problem of adaptive crowdsourcing
and develop a framework iCrowd to support adaptive crowd-
sourcing in existing crowdsourcing platforms (see Section 2).
(2) We prop ose a graph-based estimation mo d el to es-
timate th e accuracies of a worker based on her completed
tasks, which can tackle the diverse accuracies of workers
across tasks and provide accurate estimation (see Section 3).
(3) We devise an adaptive assignment framework, prove
that the optimal task assignment problem is NP-hard, and
develop a greedy algorithm to enable instant task assign-
ments (see Section 4).
(4) We deploy iCrowd on AMT and conduct extensive ex-
periments on two real datasets. Experimental results show
that iCrowd achieves 10% - 20% improvemen t on accuracy
compared with state-of-the-art approaches (see Section 6).
2. AN OVERVIEW OF ADAPTIVE CROWD-
SOURCING
2.1 Problem Statement
Microtasks. Consider a requester who publishes a set of
microtasks T = {t
1
,t
2
,...t
m
}.Foreaseofpresentation,
each microtask is a binary microtask with YES/NO choices.
Note that our techniques can be extended to microtasks with
more than two choices. Table 1 provides twelve microtasks
for entity resolution.Eachmicrotaskwantsworkerstoverify
whether two records (in the second column) are matched as
asameproductmodel. Forexample,t
1
requires workers
to verify whether iphone 4 WiFi 32GB”and“iphone four
3G black”are duplicated models. The worker, who has been
assigned with t
1
,wouldanswerYES if she agrees that they
are the same product, or NO otherwise.
Worke rs. AsetofworkersW = {w
1
,w
2
,...,w
n
} will work
on microtasks in T .Notethatworkersetincrowdsourcing
is dynamic:anyexistingworkermaybecomeinactive by
stopping work on T while new workers may become active.
Moreover, since workers are prone to errors [22], answers
provided by them may not be always correct. To predic-
twhetheraworkercancorrectlyansweramicrotask,we
introduce accuracy defined as below.
Definition 1 (Accuracy). The accuracy of a worker
w Won a microtask t
i
T,denotedbyp
w
i
,istheproba-
bility p
w
i
= Pr{w correctly answers t
i
}.
For simplicity, we use vector p
w
= {p
w
1
,p
w
2
,...,p
w
|T |
} to
represent the accuracies of w on microtasks in T .
Microtask Assignment. In crowdsourcing, to improve the
quality, a microtask is usually assigned to multiple workers
and its result is obtained via a voting scheme. Under this
scheme, we assign a microtask t
i
to a worker set W
t
i
W
with size k,wherek is an assignment size to represent the
number of w orkers that can be assigned with t
i
. k is usual-
ly provided by the crowdsourcing requester. Given worker
set W
t
i
,weutilize(weighted)majorityvoting,whichiswell
accepted in many crowdsourcing approaches [11, 32, 7]. For
ease of presentation, this paper considers the simple major-
ity voting where k is an odd number. If more than or equal
of 16
免费下载
【版权声明】本文为墨天轮用户原创内容,转载时必须标注文档的来源(墨天轮),文档链接,文档作者等基本信息,否则作者和墨天轮有权追究责任。如果您发现墨天轮中有涉嫌抄袭或者侵权的内容,欢迎发送邮件至:contact@modb.pro进行举报,并提供相关证据,一经查实,墨天轮将立刻删除相关内容。

评论

关注
最新上传
暂无内容,敬请期待...
下载排行榜
Top250 周榜 月榜