# Multiple models or one single model?

**URL:** <https://support.prodi.gy/t/multiple-models-or-one-single-model/3916>\
**Category:** Uncategorized\
**Tags:** usage, ner\
**Created:** [February 21, 2021, 4:09pm UTC](https://support.prodi.gy/t/multiple-models-or-one-single-model/3916 "2021-02-21T16:09:57Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![santoshbs](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/santoshbs/32/1954_2.png) [@santoshbs](https://support.prodi.gy/u/santoshbs)\
**Post date:** [February 21, 2021, 4:09pm UTC](https://support.prodi.gy/t/multiple-models-or-one-single-model/3916/1 "2021-02-21T16:09:58Z")

</div>

I just finished saving a model with one label (based on 600+ manual annotations of text) using the following command:

> prodigy train ner  
> my\_annotated\_data\_1  
> en\_vectors\_web\_lg  
> --init-tok2vec ../tok2vec\_cd8\_model289.bin  
> --output ./my\_model\_1  
> --eval-split 0.2`

Having obtained an `F-score` of \>95, I now would like to add a second label through the same steps of `ner.manual` and prodigy `train`.

I am not sure if should create a separate annotation dataset _my\_annotated\_data\_2_ with the second label -

- and then train and save a separate model _my\_model\_2_; or
- but train and save on the same model _my\_model\_1_ by providing both _my\_annotated\_data\_1_ and _my\_annotated\_data\_2_ as comma separated datasets to the prodigy `train` recipe

Not sure which of these is a better practice and would help achieve the most accurate results. Is there a third alternative?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 22, 2021, 1:12am UTC](https://support.prodi.gy/t/multiple-models-or-one-single-model/3916/2 "2021-02-22T01:12:52Z")

</div>

> [@santoshbs](#):
>
> but train and save on the same model _my\_model\_1_ by providing both _my\_annotated\_data\_1_ and _my\_annotated\_data\_2_ as comma separated datasets to the prodigy `train` recipe

Hi! If your goal is to have one pipeline predicting both labels, this is definitely the approach I would recommend 👆 The presence and absence of one label can always be relevant for all other labels as wel, since the entity recognizer predicts token-based tags, and named entities can't overlap.

It'll definitely be interesting to run different experiments here, though and compare the per-label evaluation scores of the joint model to models trained separately on only one label at a time. If there's a big difference here, this could point to potential problems and conflicts in the data.

A workflow you probably want to avoid is updating the same trained artifact multiple times with different datasets and different labels. This will make the process and the results much harder to reason about, and you're risking forgetting effects at every step.

---

<div class="post-metadata">

**Author:** ![santoshbs](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/santoshbs/32/1954_2.png) [@santoshbs](https://support.prodi.gy/u/santoshbs)\
**Post date:** [February 22, 2021, 8:15am UTC](https://support.prodi.gy/t/multiple-models-or-one-single-model/3916/3 "2021-02-22T08:15:37Z")

</div>

Thank you so much, @ines. I am intending to experiment with combination and separate models.
