# How do I work with available word vectors during NER training?

**URL:** <https://support.prodi.gy/t/how-do-i-work-with-available-word-vectors-during-ner-training/5754>\
**Category:** Uncategorized\
**Tags:** ner, training\
**Created:** [June 28, 2022, 2:31pm UTC](https://support.prodi.gy/t/how-do-i-work-with-available-word-vectors-during-ner-training/5754 "2022-06-28T14:31:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![nanyasrivastav](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/nanyasrivastav/32/2939_2.png) [@nanyasrivastav](https://support.prodi.gy/u/nanyasrivastav)\
**Post date:** [June 28, 2022, 2:31pm UTC](https://support.prodi.gy/t/how-do-i-work-with-available-word-vectors-during-ner-training/5754/1 "2022-06-28T14:31:01Z")

</div>

Hello,

I want to train a NER model to automatically extract "nanoparticle" entities in biomedical text. Since it is a new entity type, I decided to train the model from scratch and use existing [word vectors](http://evexdb.org/pmresources/vec-space-models/) that have been trained on PubMed corpus (biomedical text) to help lift my model's performance.

This is probably the wrong way to do it, but I downloaded the [PubMed-w2v.bin](http://evexdb.org/pmresources/vec-space-models/) file, and followed the same command used for the [food ingredients](https://github.com/explosion/projects/tree/v3/tutorials/ner_food_ingredients) example:

`python -m prodigy train ./model --ner dataset_name --base-model en_core_sci_md`  
`--paths.init-tok2vec ./PubMed-w2v.bin --eval-split 0.2`

I am not able to figure out how to incorporate the available word vectors into my workflow. The above command gives horrible results, which I guess makes sense because the word2vec file isn't exactly the same as pretrained tok2vec weights used in the food ingredients example (?)

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [June 29, 2022, 10:50am UTC](https://support.prodi.gy/t/how-do-i-work-with-available-word-vectors-during-ner-training/5754/2 "2022-06-29T10:50:03Z")

</div>

An easier way might be to initialize a new spaCy model with these vectors beforehand.

Have you seen the [init vectors](https://spacy.io/api/cli/#init-vectors) command? This will allow you to create a new spaCy model on disk that carries your embeddings in it. This local model can be then referenced as a starting point via `--base-model` in the Prodigy `train` command.

Let me know if this does not work for you, but this is how I usually use different vectors when I run benchmarks.

---

<div class="post-metadata">

**Author:** ![nanyasrivastav](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/nanyasrivastav/32/2939_2.png) [@nanyasrivastav](https://support.prodi.gy/u/nanyasrivastav)\
**Post date:** [June 29, 2022, 1:46pm UTC](https://support.prodi.gy/t/how-do-i-work-with-available-word-vectors-during-ner-training/5754/3 "2022-06-29T13:46:00Z")

</div>

Hello Vincent,

Yes, I have checked out that command. I even tried it out, but I got a bunch of errors. I guess it was because the w2v file that I have downloaded from this [website](http://evexdb.org/pmresources/vec-space-models/) is a `.bin` file, but the documentation on SpaCy's website requires it to be in the `.txt` format or a zipped text file in `.zip` or `.tar.gz` format. The size of the downloaded `.bin` file is 1.77 GB, will I have to convert it and then check if it works? I was hoping to find another way to directly use the `.bin` file using the init vectors command.

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [June 30, 2022, 6:53am UTC](https://support.prodi.gy/t/how-do-i-work-with-available-word-vectors-during-ner-training/5754/4 "2022-06-30T06:53:04Z")

</div>

Do you happen to know how these vectors were trained? With Gensim? FastText? Are you aware of any documentation for these vectors?

I'm also wondering, did you try running the other models/vectors from [scispacy](https://allenai.github.io/scispacy/)?
