# Restarting Prodigy with a new session

**URL:** <https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577>\
**Category:** Uncategorized\
**Tags:** usage, solved\
**Created:** [February 24, 2020, 7:05am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577 "2020-02-24T07:05:32Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![RohitRanga](https://avatars.discourse-cdn.com/v4/letter/r/9d8465/32.png) [@RohitRanga](https://support.prodi.gy/u/RohitRanga)\
**Post date:** [February 24, 2020, 7:05am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/1 "2020-02-24T07:05:32Z")

</div>

Hi,  
I was creating annotations manually for my dataset which is in jsonl format.  
I have a question here. Lets say I close my session and start again in a few hours. Does Prodigy make sure (in the new session) that it selects records which have not been annotated already? Thanks.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 24, 2020, 11:42am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/2 "2020-02-24T11:42:14Z")

</div>

> [@RohitRanga](#):
>
> Does Prodigy make sure (in the new session) that it selects records which have not been annotated already? Thanks.

Hi! Prodigy will skip incoming examples that are already saved in the current dataset – so if you're starting a new session with the same dataset name, you should only see examples that haven't been annotated yet.

Under the hood, Prodigy uses hashes to determine whether an incoming example is the same question, or a different question about the same data and will filter accordingly, depending on the recipe. You can read more about the mechanism here: [Loaders and Input Data · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/api-loaders#hashing)

---

<div class="post-metadata">

**Author:** ![RohitRanga](https://avatars.discourse-cdn.com/v4/letter/r/9d8465/32.png) [@RohitRanga](https://support.prodi.gy/u/RohitRanga)\
**Post date:** [February 28, 2020, 7:44am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/3 "2020-02-28T07:44:34Z")

</div>

Thanks Ines!

---

<div class="post-metadata">

**Author:** ![DiegoMartinaglia](https://avatars.discourse-cdn.com/v4/letter/d/7feea3/32.png) [@DiegoMartinaglia](https://support.prodi.gy/u/DiegoMartinaglia)\
**Post date:** [July 9, 2020, 3:39pm UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/4 "2020-07-09T15:39:37Z")

</div>

I used following command to create annotation:  
prodigy ner.teach news\_de\_v1.0 de\_core\_news\_lg ./news\_de.json

When the next day, I restarted the same work, the same questions appeared.  
What did I wrong ?

Thanks, Diego

---

<div class="post-metadata">

**Author:** ![simonschoe](https://avatars.discourse-cdn.com/v4/letter/s/ec9cab/32.png) [@simonschoe](https://support.prodi.gy/u/simonschoe)\
**Post date:** [October 21, 2022, 11:16am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/5 "2022-10-21T11:16:47Z")

</div>

@ines Is there a way to make prodigy not exclude duplicates? Let's say I have intentionally included two identical examples to later assess intra-coder agreement. In that case I would want to prevent prodigy from excluding based the `_input_has`.

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [October 24, 2022, 9:02am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/6 "2022-10-24T09:02:09Z")

</div>

While it's a bit of a hack, you could include the annotator information in the `task_hash`. This way, each annotator would be part of the task definition. This comes with some downsides, because you cannot identify each task on it's own anymore without the annotator info, but theoretically it'd do what you're asking for.

This `task_hash` is documented on our docs [here](https://prodi.gy/docs/api-loaders#hashing). In case it's of interest, the difference between the input and task hash is explained in detail here:

> [@Difference between Input hash and task hash](https://support.prodi.gy/t/difference-between-input-hash-and-task-hash/3220):
>
> I have been working with the prodigy db, and im confused as to what the difference is between \_input\_hash and \_task\_hash which are assigned to each task. I couldnt find anything satisfactory in the documentation. Can you please help?

If you're working on a custom recipe you should be able to use the [set\_hashes](https://prodi.gy/docs/api-components#set_hashes) function get the behavior. In your case, I imagine it would look something like:

```python
from prodigy import set_hashes

stream = (add_annotator_info(eg, annotator_name) for eg in stream)
stream = (set_hashes(eg, input_keys=("text",), task_keys=("label", "options", "annotator")) for eg in stream)

```

## Just be aware

That said, I do want to stress that this is a bit of a hack. If you're going to be working with multiple annotators you'd also want to have a system that can regulate the annotator overlap and this sort of functionality is planned for the Prodigy Teams product. A thread on this product, which is still in development, can be found here:

> [@sparkles Prodigy Annotation Manager Update // Prodigy Scale // Prodigy Teams](https://support.prodi.gy/t/prodigy-annotation-manager-update-prodigy-scale-prodigy-teams/805):
>
> Many of you have been asking about how to scale your Prodigy projects – including how to manage more annotators, how to keep multiple feeds running, and how to make sure your annotators stay in agreement. These are challenging problems that are outside of the scope of the [Prodigy](https://prodi.gy) software itself, so we've always planned an extension product that helps you use Prodigy in larger products. We've been quietly working away on this for some time, so we're excited to finally announce our progress! The …

---

<div class="post-metadata">

**Author:** ![simonschoe](https://avatars.discourse-cdn.com/v4/letter/s/ec9cab/32.png) [@simonschoe](https://support.prodi.gy/u/simonschoe)\
**Post date:** [October 25, 2022, 6:04am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/7 "2022-10-25T06:04:34Z")

</div>

@koaning thanks for the considerate reply Vincent! Currently, I have a single prodigy server running per annotator, so mediating between different annotators during labeling is not an issue.

I now rely on the following workaround (similar to your proposal):

1. Add an additional field to the source `.jsonl` file (called `DUPL`) that indicates whether it is the first or second appearance of a given sentence. (could have maybe also done it within the recipe using a function similar to `add_annotator_info`)
2. Use this additional meta field to compute a new, unique hash per example:

```python
stream = JSONL(source)
stream = (prodigy.set_hashes(eg, input_keys=('text', 'DUPL')) for eg in stream)

```

The way I see it, this solves the issue. Maybe two quick follow-ups:

1. Is there any argument for integrating the additional meta field in the `input_keys` vs. `task_keys` hash?
2. Isn't the assessment of intra-coder reliability a common use case? Just wondering, bc it seems like this is a somewhat hacky workaround that could be more easily integrated into prodigy.

Thanks for taking the time and helping out!

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [October 26, 2022, 9:21am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/8 "2022-10-26T09:21:37Z")

</div>

Two short answers!

### Question 1

One idea behind having`input_keys` and `task_keys` seperately is so that you can always add a new label later. You can, for example, start with two classes in a classification problem but always add a 3rd one easily. If the `input_keys` were merged with the `task_keys` you wouldn't be able to do that.

Hopefully, this argument also paints a picture of warning. While nothing is stopping you from adding whatever info you like, you should try and think about future changes that might be impacted. If the metadata really makes it a new task, or a new training example, then you can consider adding it. If not, you risk loosing the ability to make each example unique later.

### Question 2

Annotator agreement is indeed a common theme/problem in our space. If you're comfortable with Python you can always implement your own solution but features surrounding agreement are planned for Prodigy Teams.

If you'd like to write your own Python solution, you'd might get away with a `groupby(input_key, task_key)` to find instances of annotator disagreement. This assumes however that the labels/task never change.

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [October 26, 2022, 9:32am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/9 "2022-10-26T09:32:12Z")

</div>

Actually! Just to check; are you aware of the `session` mechanic? The one explained here:

> [@prodigy Multi-user session access](https://support.prodi.gy/t/prodigy-multi-user-session-access/1644/2):
>
> Hi! If you're looking for user management and monitoring at that scale, this is probably outside of the scope of Prodigy standalone. You might want to have a look at the upcoming Prodigy Scale, which is a separate product we're developing for exactly that use case. See here for details: While Prodigy includes some features you can use to build a multi-user system, you still have to decide how you want it to be set up. For instance, how do you want the authentication to work? How should the us…

Here's what `db-out` would look like for a simple text classification usecase.

```python
{"text":"this is a single example yo","_input_hash":-465404500,"_task_hash":221834242,"label":"demo","_view_id":"classification","answer":"accept","_timestamp":1666776571,"_annotator_id":"issue-6042-foobar","_session_id":"issue-6042-foobar"}
{"text":"this is a single example yo","_input_hash":-465404500,"_task_hash":221834242,"label":"demo","_view_id":"classification","answer":"accept","_timestamp":1666776577,"_annotator_id":"issue-6042-vincent","_session_id":"issue-6042-vincent"}

```

That should give you access to `_annotator_id` as well. Isn't that what you'd want?

---

<div class="post-metadata">

**Author:** ![simonschoe](https://avatars.discourse-cdn.com/v4/letter/s/ec9cab/32.png) [@simonschoe](https://support.prodi.gy/u/simonschoe)\
**Post date:** [October 28, 2022, 7:24am UTC](https://support.prodi.gy/t/restarting-prodigy-with-a-new-session/2577/10 "2022-10-28T07:24:53Z")

</div>

> That should give you access to `_annotator_id` as well. Isn't that what you'd want?

Currently, we are running individual prodigy servers for each annotator, which is why the `_annotator_id` never occurs. Actually, I think I'll try reimplementing the workflow using multi-user sessions which appears way more elegant.

However, my initial concern/issue was simply about intra-coder (instead of inter-coder) agreement, that is, how consistent a single annotator is over time; thats why I was trying to mix in duplicate inputs (for the same annotator) to assess that consistency. And I believe `set_hashes` offered me a nice and clean solution for that by simply mixing in metadata about whether or not that is the first, second, or third appearance of a given sentence. That way, the same sentence wasn't filtered out by prodigy.

Long story short: The problem is solved! 🤗 Thanks for providing context and helping out along the way!
