# Which base\_model to use for ner.manual

**URL:** <https://support.prodi.gy/t/which-base-model-to-use-for-ner-manual/7547>\
**Category:** Uncategorized\
**Tags:** ner, spacy\
**Created:** [April 27, 2025, 1:21am UTC](https://support.prodi.gy/t/which-base-model-to-use-for-ner-manual/7547 "2025-04-27T01:21:49Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [April 28, 2025, 9:20am UTC](https://support.prodi.gy/t/which-base-model-to-use-for-ner-manual/7547/2 "2025-04-28T09:20:27Z")

</div>

Hi @Fangjian,

I understand you want to train an NER model to recognize custom entities i.e. entities that are not covered by `en_core_scibert`? In that case, it is indeed better to train your NER model from scratch. Trying to add new categories to already trained model might result in unpredictable behavior as it's hard to control how the new data affects already existing weights. Especially that the pretrained categories were probably trained on a much bigger dataset. In this[post](https://support.prodi.gy/t/work-flow-for-extending-an-ner-model-with-new-entity-types/1603/2) Matt explains the dynamics of resuming the training of a NER component.

So in your case, you want to substitute the `en_core_scibert` NER component with your custom (blank) NER component. You can check this [example spaCy project](https://github.com/explosion/projects/tree/v3/pipelines/ner_demo_replace) to see how it can be done.  
I'm not sure if there are other components of `en_core_scibert` that depend on NER predictions. If that's the case you might need to remove them as well or leave everything as is and give your NER component a different name and freeze the `en_core_sci_bert` during training. You can read a bit more about customizing spaCy pipelines [here](https://spacy.io/usage/training#config-custom).

Even though you'd be training your NER component from scratch, you can still benefit from `en_core_scibert` tokenizer and word vectors and/or pretrained embeddings both for the training and the data annotation with `ner.manual`. In other words, you can benefit from transfer learning at the representation level independent of the pre-trained NER component (that you'll discard). In fact, it's very important that the same tokenizer is used during the data annotation, model training and later in production. So, in summary, your workflow should look like this:

1. annotate using `ner.manual` and `en_core_scibert` as the base model to benefit from the scientific tokenizer (I'm assuming the labels you'd be using will be different from the `en_core_scibert` pretrained ones)
2. write a spaCy training config where you substitute the pretrained NER with your custom NER (as in [this example project](https://github.com/explosion/projects/tree/v3/pipelines/ner_demo_replace))

---

_[View the full topic](https://support.prodi.gy/t/which-base-model-to-use-for-ner-manual/7547)._
