# add new lables as per new data received to existing data set and retrain the NER model

**URL:** <https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885>\
**Category:** Uncategorized\
**Tags:** ner, spacy\
**Created:** [August 26, 2022, 12:48pm UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885 "2022-08-26T12:48:18Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vishal112](https://avatars.discourse-cdn.com/v4/letter/v/7feea3/32.png) [@Vishal112](https://support.prodi.gy/u/Vishal112)\
**Post date:** [August 26, 2022, 12:48pm UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/1 "2022-08-26T12:48:18Z")

</div>

Pls guide me regarding this

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [August 26, 2022, 3:15pm UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/2 "2022-08-26T15:15:33Z")

</div>

hi @Vishal112!

That's a bit of a tricky question.

One of the first questions you should ask isn't really a ML question but a business question: what's the expectation on the frequency/timing for new NER labels? and when changed, how are annotation guidelines updated so that your annotators have clear definitions of the new labels?

If you have a model in production, I would caution against an expectation that you can add/retrain many times and on an ad hoc (non-regular occuring) basis (it would be okay if you're only in model development developing the model though). The reason is it may be very difficult to track changes so you likely should agree with your model users (stakeholders) on fixed times to add new labels (e.g., once a month).

Along those same lines, what's important (and sometimes forgotten) is to ensure you have a clear definition of what you're labeling through explicit annotation guidelines. This can be as simple as definitions of your entities. One great example is from the Guardian, who published an article about how they used Prodigy for a `ner` model:

> **[Talking sense: using machine learning to understand quotes](https://www.theguardian.com/info/2021/nov/25/talking-sense-using-machine-learning-to-understand-quotes)**
>
> The Guardian’s data scientists have been working with other newsrooms on a global project to think about AI and journalism. Here they explain how they have been teaching a machine to understand what a quote is

They published their code and [their annotation guidelines](https://github.com/JournalismAI-2021-Quotes/quote-extraction/blob/main/annotation_rules/Quote%20annotation%20guide.pdf). As they discussed, it's important to have routine discussions with your modelers and business stakeholders to constantly update those guidelines.

The best example of this philosophy is in Matt's 2018 talk where he talks about successful ways of defining the business problem, the need for clear annotation guidelines, and role the that an iterative approach can aid in resulting in successful vs. unsuccessful ML projects:

[![](https://img.youtube.com/vi/jpWqz85F_4Y/maxresdefault.jpg "Building new NLP solutions with spaCy and Prodigy - Matthew Honnibal") ](https://www.youtube.com/watch?v=jpWqz85F_4Y&t=372)

Now I suspect you're more interested in how to implement updating NER labels and retrain the model, there are a lot of past Prodigy Support issues and documentation that can help.

- First, I'd start with our [NER workflow](https://prodi.gy/prodigy_flowchart_ner-36f76cffd9cb4ef653a21ee78659d366.pdf). We're hoping very soon to update this as some of the Prodigy recipe names have changed (e.g., `train` now instead of `batch.train`). But the main idea still holds. An important part is whether you're using a pre-trained model (e.g., `en_core_web_sm`) to update and/or add new existing entity types or starting a new model. What's important is that we recommend if you're adding more than 3+ new entity types, you're likely better off starting from scratch.

- If you have some prior knowledge about your new entity (e.g., terms that are related), you should also consider adding [match patterns](https://prodi.gy/docs/named-entity-recognition#patterns) to help bootstrap your entities. This will make it easier to label as these patterns will come highlighted.

- If you are starting from scratch and simply want to create a workflow for a few new entities, I like my teammate's @ljvmiranda921 advice in this post that is very similar to your question:

> [@Adding new label](https://support.prodi.gy/t/adding-new-label/4861/2):
>
> Hi @pkras! Is there a way to use ner.manual or ner.correct on the saved labelled dataset and add the new label? It's possible to [load existing datasets again](https://prodi.gy/docs/api-loaders#datasets). Prodigy will load the anotations from the dataset then stream them back in. Perhaps you can try that then include your new label: prodigy ner.manual \<new\_dataset\> \<model\> dataset:\<old\_dataset\> --label PERSON,ORG Similarly, i would like to ask what you would advise in the case of wanting to remove a label from the existing scheme. D…

- It's also good to be aware that there are ways to create "nested" NER labels:

> [@Nested labels for NER](https://support.prodi.gy/t/nested-labels-for-ner/696):
>
> Is it possible to create a nested label structure for NER? For example, I have a new entity - DRUG, but also there are several subtypes: ‘antidepressants’, ‘sedative’, ‘cardiac’ etc. Something like: drug\_list.jsonl -\> {“label”:“DRUG”,“subtype”:“antidepressant”, “pattern”:[{“lower”:“citalopram”}]} {“label”:“DRUG”, “subtype”:“sedative”, “pattern”:[{“lower”:“clonazepam”}]} Or any other solution how to have access to entities and their subcategories. Thanks.

- Last, it's important to know that Prodigy wasn't designed to create "labels-on-the-fly" and that label sets are fixed for each annotation session. You can definitely add new labels in between annotation sessions, but I want to make sure you understand that Prodigy isn't designed to do this within annotation session.

> [@NER - Add labels on the fly](https://support.prodi.gy/t/ner-add-labels-on-the-fly/4206/2):
>
> Hi! Prodigy expects you to define the label scheme when you start the annotation process, and if your goal is to collect annotations for machine learning, you typically do not want the annotator to be able to decide about your label scheme and enter labels manually. The presence and absence of a given label is very important and will have a big impact on your entire model. Also, if a new label is introduced later, this can potentially invalidate previous annotations and you'll end up with incon…

Be sure to continue searching on Prodigy support. There may be several other posts -- I found these after a quick search.

I hope this helps and let us know if you have further questions!

---

<div class="post-metadata">

**Author:** ![Vishal112](https://avatars.discourse-cdn.com/v4/letter/v/7feea3/32.png) [@Vishal112](https://support.prodi.gy/u/Vishal112)\
**Post date:** [August 29, 2022, 4:26am UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/3 "2022-08-29T04:26:28Z")

</div>

I have new categories to train and don't want to train the model from scratch, currently I am testing my model in local environment and the data that I have annotated is very large hence I want to skip this step and want to add new lables for new categories I received

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [August 29, 2022, 1:10pm UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/4 "2022-08-29T13:10:57Z")

</div>

hi @Vishal112!

> [@Vishal112](#):
>
> the data that I have annotated is very large hence I want to skip this step and want to add new lables for new categories I received

If you decide to retrain (i.e., don't start with a `blank:en` but a previously trained model), be aware, you'll likely run into problems of catastrophic forgetting. There's not an easy solution to this except being aware that if you retrain with an imbalance of examples (e.g., only retraining on new NER entities), your model may forget the old entities.

> [@New entity model ruins other entities](https://support.prodi.gy/t/new-entity-model-ruins-other-entities/179/2):
>
> Sorry about the late reply! I think what you’re experiencing might be whats often referred to as the “catastrophic forgetting problem”. As your model is learning about the new entity type, it’s “forgetting” what it has previously learned. In your example, this is pretty significant – but it might be because you’ve trained a completely new entity, so the only data the model is updating on is examples labelled TECH and none of the other entity types. Because the model is never “reminded” about the…

---

<div class="post-metadata">

**Author:** ![Vishal112](https://avatars.discourse-cdn.com/v4/letter/v/7feea3/32.png) [@Vishal112](https://support.prodi.gy/u/Vishal112)\
**Post date:** [September 5, 2022, 7:39am UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/5 "2022-09-05T07:39:12Z")

</div>

no this is not my concern @ines

consider my scenario  
I have 2000 email data with subject(category/Lable) I have annotated the data manually ner.manual receipe  
then I have trained the model and I am using it

now I got another 1000 email data with different lables or categories now I dont want to annotate that data which I have annotated previously so what I supposed to do now

pls guide

---

<div class="post-metadata">

**Author:** ![Vishal112](https://avatars.discourse-cdn.com/v4/letter/v/7feea3/32.png) [@Vishal112](https://support.prodi.gy/u/Vishal112)\
**Post date:** [September 6, 2022, 9:50am UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/6 "2022-09-06T09:50:19Z")

</div>

one more thing pls address this issue too why I am getting all label together  
 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/e/e3c7442d306f20bff304fb56944ffeca0cc3d5ce.png)

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [September 6, 2022, 6:44pm UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/7 "2022-09-06T18:44:09Z")

</div>

Hi @Vishal112,

> [@Vishal112](#):
>
> consider my scenario  
> I have 2000 email data with subject(category/Lable) I have annotated the data manually ner.manual receipe  
> then I have trained the model and I am using it
> 
> now I got another 1000 email data with different lables or categories now I dont want to annotate that data which I have annotated previously so what I supposed to do now

If your new annotation round has different labels, why are you trying to use the model you previously created? Why not create two separate models? Is there overlap in the labels? If so in what way?

> [@Vishal112](#):
>
> one more thing pls address this issue too why I am getting all label together

I suspect there's an issue with your patterns. Can you provide what your patterns look like? That span is hitting your 3 match pattern (see bottom right) -- you likely are having issues with that pattern.

> [@Vishal112](#):
>
> ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/e/e3c7442d306f20bff304fb56944ffeca0cc3d5ce.png)

Is your ultimate goal an intent model for chats?

If so, why not use a text classification UI instead of a NER for your intent annotations?

Most intent-models do have an accompanying NER model but those NER provide context/additional to help take an action on the intent. You seem to be using an NER for intent (i.e., whether this is account-related, credit-card-related, etc.).

---

<div class="post-metadata">

**Author:** ![Vishal112](https://avatars.discourse-cdn.com/v4/letter/v/7feea3/32.png) [@Vishal112](https://support.prodi.gy/u/Vishal112)\
**Post date:** [September 7, 2022, 10:00am UTC](https://support.prodi.gy/t/add-new-lables-as-per-new-data-received-to-existing-data-set-and-retrain-the-ner-model/5885/8 "2022-09-07T10:00:11Z")

</div>

No data is same of email bodies but categories are different
