# ner.teach does not exclude dataset even after '--exclude'

**URL:** <https://support.prodi.gy/t/ner-teach-does-not-exclude-dataset-even-after-exclude/1192>\
**Category:** Uncategorized\
**Tags:** usage, ner\
**Created:** [February 5, 2019, 11:24pm UTC](https://support.prodi.gy/t/ner-teach-does-not-exclude-dataset-even-after-exclude/1192 "2019-02-05T23:24:36Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Arul](https://avatars.discourse-cdn.com/v4/letter/a/b2d939/32.png) [@Arul](https://support.prodi.gy/u/Arul)\
**Post date:** [February 5, 2019, 11:24pm UTC](https://support.prodi.gy/t/ner-teach-does-not-exclude-dataset-even-after-exclude/1192/1 "2019-02-05T23:24:37Z")

</div>

Hi,  
I am trying to label another round of data with the existing training dataset using ner.teach. I already have one set annotated in dataset “training\_1” (silver). My input file has a lot of text data in csv which was used as input for “training\_1” (a part of it was done in first round). Now, when i use this command with these args, prodigy should consider the text that is not in ‘training\_1’. But in the interface, i am getting the text that was already labeled in ‘training\_1’ dataset.

```
prodigy ner.teach training_2 trained_models_spacy long_text_train.csv --label Labels.txt --patterns Prodigy_Patterns.jsonl --exclude training_1

```

I dont know why this is not working. Should i match and drop the already tagged text before i give input as csv?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 6, 2019, 10:33am UTC](https://support.prodi.gy/t/ner-teach-does-not-exclude-dataset-even-after-exclude/1192/2 "2019-02-06T10:33:08Z")

</div>

Hi! Are you sure the examples you’re seeing are actually the _same_ examples? So, the same span suggestion on the same text? The thing is, the exclude mechanism will only look at _identical_ examples to ensure that you’re never annotating the same question twice. But if there’s a different question on the same text – for example, with a different entity span suggested – you’ll still get to see that example, because it’s a different question.

If you don’t want to include examples with _texts_ you’ve already annotated something on, you could write your own stream filter that gets the input hashes from the dataset and only sends out examples with a different input hash. For details, check out the docs on the `filter_inputs` helper in your `PRODIGY_README.html`.

---

<div class="post-metadata">

**Author:** ![Arul](https://avatars.discourse-cdn.com/v4/letter/a/b2d939/32.png) [@Arul](https://support.prodi.gy/u/Arul)\
**Post date:** [February 6, 2019, 6:43pm UTC](https://support.prodi.gy/t/ner-teach-does-not-exclude-dataset-even-after-exclude/1192/3 "2019-02-06T18:43:50Z")

</div>

Ohh, ok. The entity suggestions now are different. But my original corpus (from which i got the model in the loop) is manually corrected using ner.teach before. Now i am using the improved model in the loop. Do you suggest adding more corpus that excludes these texts or do you suggest looping again with the same text with the correcting different labels suggested?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 6, 2019, 6:48pm UTC](https://support.prodi.gy/t/ner-teach-does-not-exclude-dataset-even-after-exclude/1192/4 "2019-02-06T18:48:21Z")

</div>

I guess it depends on how much raw data you have! If you have a lot more raw text, then yes, maybe you can try showing it something else instead of looping over the same data again. If you only have limited examples, then I’d say it’s okay to start at the beginning again. When you first annotated the data, you also didn’t get to see _every_ example. The active learning will skip examples and only show you the most relevant ones. When you loop over it again with a different model, you might also see different suggestions.

---

<div class="post-metadata">

**Author:** ![Arul](https://avatars.discourse-cdn.com/v4/letter/a/b2d939/32.png) [@Arul](https://support.prodi.gy/u/Arul)\
**Post date:** [February 6, 2019, 6:54pm UTC](https://support.prodi.gy/t/ner-teach-does-not-exclude-dataset-even-after-exclude/1192/5 "2019-02-06T18:54:50Z")

</div>

Thank you. I have a good amount of raw data. I will exclude these texts and give it new text.
