# Iterating on a NER spaCy model with Prodigy

**URL:** <https://support.prodi.gy/t/iterating-on-a-ner-spacy-model-with-prodigy/3201>\
**Category:** Uncategorized\
**Tags:** usage, ner, spacy, solved\
**Created:** [July 17, 2020, 10:35pm UTC](https://support.prodi.gy/t/iterating-on-a-ner-spacy-model-with-prodigy/3201 "2020-07-17T22:35:34Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![meatball\_nlp](https://avatars.discourse-cdn.com/v4/letter/m/0ea827/32.png) [@meatball\_nlp](https://support.prodi.gy/u/meatball_nlp)\
**Post date:** [July 17, 2020, 10:35pm UTC](https://support.prodi.gy/t/iterating-on-a-ner-spacy-model-with-prodigy/3201/1 "2020-07-17T22:35:34Z")

</div>

Hello,

I have a general question about Prodigy best practices. I bootstrapped a spaCy model using raw data and `ner.manual` command to train a model on top of `en_core_web_md` to detect a custom entity.

I evaluated my model and found some errors, so I collected more data and now want to update the old model with this new data. What is the best practice here? Should I use `ner.correct` or `ner.manual` - is the only difference that one suggests annotations to speed up time?

Also before training with `prodigy train ner` should I combine the two datasets into one to prevent anything like the catastrophic forgetting problem? Thanks!

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 18, 2020, 10:25am UTC](https://support.prodi.gy/t/iterating-on-a-ner-spacy-model-with-prodigy/3201/2 "2020-07-18T10:25:47Z")

</div>

> [@meatball\_nlp](#):
>
> What is the best practice here? Should I use `ner.correct` or `ner.manual` - is the only difference that one suggests annotations to speed up time?

Yes, exactly – and it makes it easy to include everything else that the model previously predicted. So if you have a model that predicts something relevant, using `ner.correct` makes more sense IMO because it saves you time _and_ lets you review the model's predictions at the same time.

> [@meatball\_nlp](#):
>
> Also before training with `prodigy train ner` should I combine the two datasets into one to prevent anything like the catastrophic forgetting problem? Thanks!

Yes, if you have all annotations, it makes sense to train from scratch / from the base model every time instead of "incrementally" updating the same artifact. It helps prevent catastrophic forgetting and other side-effects, and also gives you a more stable evaluation. (If you have a separate evaluation dataset, you can always evaluate the same process on the same data but you couldn't really compare the results unless you always start with the same base model.)

The `train` command lets you specify a comma-separate list of datasets btw and will take care of merging the annotations. So you can run: `prodigy train ner dataset_one,dataset_two ...`

---

<div class="post-metadata">

**Author:** ![meatball\_nlp](https://avatars.discourse-cdn.com/v4/letter/m/0ea827/32.png) [@meatball\_nlp](https://support.prodi.gy/u/meatball_nlp)\
**Post date:** [July 20, 2020, 5:25pm UTC](https://support.prodi.gy/t/iterating-on-a-ner-spacy-model-with-prodigy/3201/3 "2020-07-20T17:25:28Z")

</div>

Awesome @ines - thanks for clarifying.

So as a follow up question when you say `it makes sense to train from scratch / from the base model every time instead of "incrementally" updating the same artifact.`

If my model is based off of `en_core_web_md`, I would always want to train off of this base model.

Your suggestion:

1. Train `en_core_web_md` with entity type X
2. Collect more data, Train `en_core_web_md` again with entity type X and entity type Y

Vs. my original thought:

1. Train `en_core_web_md` with entity type X -\> `en_core_web_md_with_x`
2. Collect more data and train `en_core_web_md_with_x` with entity type X and entity type Y

At what point would I not want to always train from scratch - when it becomes too slow?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 21, 2020, 9:04am UTC](https://support.prodi.gy/t/iterating-on-a-ner-spacy-model-with-prodigy/3201/4 "2020-07-21T09:04:44Z")

</div>

> [@meatball\_nlp](#):
>
> Your suggestion:
> 
> 1. Train `en_core_web_md` with entity type X
> 2. Collect more data, Train `en_core_web_md` again with entity type X and entity type Y

Yes, exactly! 👍

> [@meatball\_nlp](#):
>
> At what point would I not want to always train from scratch - when it becomes too slow?

That shouldn't be a problem – even if it's slower, it would normaly be worth it. If you have access to all the data, it makes sense to train with the whole corpus, from scratch.

Starting with an existing model makes sense if you don't have access to the whole original training data or training process – like with the `en_core_web_sm` model. We can distribute the model, but not the data, and if you don't have a license for it, you couldn't just train it from scratch. So if you just want to adjust its predictions on your data, it can be easier and more efficient to just take the existing model and update it with a few more examples.
