# Help with training from scratch english NER model with pretrained Gensim vectors

**URL:** <https://support.prodi.gy/t/help-with-training-from-scratch-english-ner-model-with-pretrained-gensim-vectors/5215>\
**Category:** Uncategorized\
**Tags:** usage, ner, spacy\
**Created:** [January 19, 2022, 9:25pm UTC](https://support.prodi.gy/t/help-with-training-from-scratch-english-ner-model-with-pretrained-gensim-vectors/5215 "2022-01-19T21:25:37Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jhutton1121](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@Jhutton1121](https://support.prodi.gy/u/Jhutton1121)\
**Post date:** [January 19, 2022, 9:25pm UTC](https://support.prodi.gy/t/help-with-training-from-scratch-english-ner-model-with-pretrained-gensim-vectors/5215/1 "2022-01-19T21:25:37Z")

</div>

I am trying to train a from scratch NER model with custom labels. I have word vectors that are pretrained from Gensim on a large corpus of r/wallstreetbets data. I need help determining the workflow from here to creating a preliminary model.

Right now I have the vectors as well as a labeled dataset containing 270 examples in my database. I'd like to train a model using my word vectors that will then be used with the ner.correct recipe.

Any help on the steps to do this is appreciated.

---

<div class="post-metadata">

**Author:** ![adriane](https://avatars.discourse-cdn.com/v4/letter/a/46a35a/32.png) [@adriane](https://support.prodi.gy/u/adriane)\
**Post date:** [January 27, 2022, 8:56am UTC](https://support.prodi.gy/t/help-with-training-from-scratch-english-ner-model-with-pretrained-gensim-vectors/5215/2 "2022-01-27T08:56:40Z")

</div>

Hi! If you have your vectors exported in word2vec text format from gensim (`save_word2vec_format`), you can initialize a base model with `spacy init vectors`:

```shell
python -m spacy init vectors en /path/to/vectors.vec /path/to/spacy_vectors

```

Then use `/path/to/spacy_vectors` as the base model when training with prodigy:

prodigy docs (for `--base-model`): [https://prodi.gy/docs/recipes#training](https://prodi.gy/docs/recipes#training)

spacy docs (for `spacy init vectors`): [https://spacy.io/usage/linguistic-features#adding-vectors](https://spacy.io/usage/linguistic-features#adding-vectors)

---

<div class="post-metadata">

**Author:** ![adriane](https://avatars.discourse-cdn.com/v4/letter/a/46a35a/32.png) [@adriane](https://support.prodi.gy/u/adriane)\
**Post date:** [January 27, 2022, 9:04am UTC](https://support.prodi.gy/t/help-with-training-from-scratch-english-ner-model-with-pretrained-gensim-vectors/5215/3 "2022-01-27T09:04:08Z")

</div>

Ah, wait, I was wrong about the prodigy side of things. Only using `--base-model` doesn't actually enable the vectors in the new `ner` component while training in prodigy by default. Let me have a look...

Edited to add:

One option is to generate a config with vectors using `spacy init config -o accuracy` and the set the vectors location in a `prodigy train` override:

```shell
spacy init config -l en -p ner -o accuracy /path/to/config.cfg
prodigy train --ner dataset --config /path/to/config.cfg --initialize.vectors /path/to/spacy_vectors

```

This should be easier, though...
