# Multiple users review same image dataset

**URL:** <https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800>\
**Category:** Uncategorized\
**Tags:** usage, image, solved\
**Created:** [September 6, 2018, 4:00am UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800 "2018-09-06T04:00:47Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![haoxi911](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/haoxi911/32/431_2.png) [@haoxi911](https://support.prodi.gy/u/haoxi911)\
**Post date:** [September 6, 2018, 4:00am UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800/1 "2018-09-06T04:00:47Z")

</div>

So we have 29k images that need to be reviewed, we installed Prodigy on a Ubuntu server, and give the link to a few teammates. They will probably review this dataset at the same time.

The question is: if more than one users open a session, they will see the same images from start to finish in their own session? if two people review the same images and one accepts and one rejects, the last one wins?

Can you please clarify?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [September 6, 2018, 12:30pm UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800/2 "2018-09-06T12:30:43Z")

</div>

> [@haoxi911](#):
>
> The question is: if more than one users open a session, they will see the same images from start to finish in their own session?

The way the web app works is actually pretty straightforward: Everytime a user opens the site, it makes a request to the `/get_questions` endpoint, which will return the next batch from the stream. This means that if two clients connect to the same session, they will get different batches of data, whatever is next up on the queue.

If you want to have multiple people annotate the same data, I'd recommend starting multiple instances – for example, run Prodigy on different ports (e.g. by setting the `PRODIGY_PORT` environment variable when you execute the command). Each annotator could then also have their own dedicated dataset that their answers are saved to. This means you'll be able to compare the work performed by the individual people.

> [@haoxi911](#):
>
> if two people review the same images and one accepts and one rejects, the last one wins?

There's no simple answer for this and how you want to use conflicting annotations later on is something you have to decide. If annotators all add to their own datasets, you'll be able to export the data and compare it to find and resolve conflicts.

One strategy could be to take the datasets, find answers with the same `_task_hash` (same question) but with different answers. You could then use a threshold of, say, 80% agreement to decide whether to include the example or not. So if 80% of annotators agree, you include the example – otherwise, you don't, or reannotate it yourself to make the final decision.

Maybe you'll also find that it's usually the same annotator who disagrees with everyone else – this could indicate a misunderstanding about the annotation scheme. This is obviously super important and something you want to find out as soon as possible. So I'd recommend exporting and analysing the data with this type of objective very early on in the process.

I'd also recommend checking out the following addon, which was developed by a fellow Prodigy user. It includes a range of features to use the tool with multiple annotators and get stats and analytics:

> **[GitHub - ahalterman/multiuser\_prodigy: Running Prodigy for a team of annotators](https://github.com/ahalterman/multiuser_prodigy)**
>
> Running Prodigy for a team of annotators. Contribute to ahalterman/multiuser\_prodigy development by creating an account on GitHub.

We're also working on an extension product, the Prodigy Annotation Manager, which is very close to a public beta now 🎉 The app will have a service component and let you manage multiple users, analyse their results and performance, create complex annotation workflows interactively and build larger corpora and labelled datasets. If that's sounds relevant, definitely keep an eye on the forum for the official announcement.

---

<div class="post-metadata">

**Author:** ![haoxi911](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/haoxi911/32/431_2.png) [@haoxi911](https://support.prodi.gy/u/haoxi911)\
**Post date:** [September 7, 2018, 2:26am UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800/3 "2018-09-07T02:26:19Z")

</div>

Thanks for your explanation, it does make sense.

One additional question (irrelevant to this one). So I am the only person who review the images, I started Prodigy and reviewed about 200 images, then I stopped Prodigy.

After about 2 hours, I restarted Prodigy and tried to continue and review more. And I realized that Prodigy loaded some images that I have reviewed hours ago.

I can see that the data were stored in sqlite database correctly, but why Prodigy didn’t load the reviewed records when we restart it? Any configurations required?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [September 7, 2018, 8:19am UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800/4 "2018-09-07T08:19:04Z")

</div>

> [@haoxi911](#):
>
> I can see that the data were stored in sqlite database correctly, but why Prodigy didn’t load the reviewed records when we restart it? Any configurations required?

Yes, by default, Prodigy makes no assumptions what the current dataset "means". But you can use the `--exclude` option to explicitly tell it to exclude examples present in one or more datasets – for example: `--exclude dataset_one,dataset_two`.

---

<div class="post-metadata">

**Author:** ![haoxi911](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/haoxi911/32/431_2.png) [@haoxi911](https://support.prodi.gy/u/haoxi911)\
**Post date:** [September 9, 2018, 3:08pm UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800/5 "2018-09-09T15:08:46Z")

</div>

Thank you, it works. Just to make sure, I can exclude the same dataset which I currently used, this way, it will pick up where it left of and continue the rest of images?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [September 10, 2018, 9:45am UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800/6 "2018-09-10T09:45:05Z")

</div>

> [@haoxi911](#):
>
> Just to make sure, I can exclude the same dataset which I currently used, this way, it will pick up where it left of and continue the rest of images?

Yes, exactly. Or, more precisely, Prodigy will skip an incoming example if it's already present in the dataset (i.e. if it has the same `_task_hash` property).

---

<div class="post-metadata">

**Author:** ![haoxi911](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/haoxi911/32/431_2.png) [@haoxi911](https://support.prodi.gy/u/haoxi911)\
**Post date:** [September 10, 2018, 9:46am UTC](https://support.prodi.gy/t/multiple-users-review-same-image-dataset/800/7 "2018-09-10T09:46:21Z")

</div>

Thank you!
