# biomedical nlp models in spacy

**URL:** <https://support.prodi.gy/t/biomedical-nlp-models-in-spacy/319>\
**Category:** Uncategorized\
**Tags:** usage, spacy, solved, gensim\
**Created:** [February 18, 2018, 12:31pm UTC](https://support.prodi.gy/t/biomedical-nlp-models-in-spacy/319 "2018-02-18T12:31:18Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![madhujahagirdar](https://avatars.discourse-cdn.com/v4/letter/m/e99b99/32.png) [@madhujahagirdar](https://support.prodi.gy/u/madhujahagirdar)\
**Post date:** [February 18, 2018, 12:31pm UTC](https://support.prodi.gy/t/biomedical-nlp-models-in-spacy/319/1 "2018-02-18T12:31:18Z")

</div>

[http://evexdb.org/pmresources/vec-space-models/](http://evexdb.org/pmresources/vec-space-models/)

We have great resources of word2vec model on biomedical text generated using gensim. Can we load them in spacy like any other model pointing to a directory ?

---

<div class="post-metadata">

**Author:** ![honnibal](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/honnibal/32/35_2.png) [@honnibal](https://support.prodi.gy/u/honnibal)\
**Post date:** [February 18, 2018, 5:53pm UTC](https://support.prodi.gy/t/biomedical-nlp-models-in-spacy/319/2 "2018-02-18T17:53:32Z")

</div>

Yes, you’ll be able to load these vectors with spaCy and use them with Prodigy. You can either create load the vectors in a custom recipe, or create a script that loads the vectors in a spaCy model and then saves the model to a directory, with `nlp.to_disk()`. Once the model has been saved to a directory, the vectors should be there, ready to use.

The only thing to keep in mind is, if you’re loading in your own vectors, you should base your model on `en_core_web_sm`. Both `en_core_web_md` and `en_core_web_lg` use pre-trained vectors as features in the tagger, parser and NER models. This means that if you replace the built-in vectors with other vectors in those models, you’ll mess up the predictions.

---

<div class="post-metadata">

**Author:** ![madhujahagirdar](https://avatars.discourse-cdn.com/v4/letter/m/e99b99/32.png) [@madhujahagirdar](https://support.prodi.gy/u/madhujahagirdar)\
**Post date:** [February 28, 2018, 2:10pm UTC](https://support.prodi.gy/t/biomedical-nlp-models-in-spacy/319/3 "2018-02-28T14:10:27Z")

</div>

I first converted the word2vec file to txt using gensim like below:

model = KeyedVectors.load\_word2vec\_format(’/Users/philips/Downloads/wikipedia-pubmed-and-PMC-w2v.bin’, binary=True)  
model.wv.save\_word2vec\_format(’/Users/philips/Downloads/wikipedia-pubmed-and-PMC-w2v.txt’)

and then

I have used the following script to save vector to disk and used language as “en”. Does that sound right?

> <https://github.com/explosion/spacy/blob/master/examples/vectors_fast_text.py>

---

<div class="post-metadata">

**Author:** ![madhujahagirdar](https://avatars.discourse-cdn.com/v4/letter/m/e99b99/32.png) [@madhujahagirdar](https://support.prodi.gy/u/madhujahagirdar)\
**Post date:** [February 28, 2018, 2:11pm UTC](https://support.prodi.gy/t/biomedical-nlp-models-in-spacy/319/4 "2018-02-28T14:11:29Z")

</div>

> [@honnibal](#):
>
> The only thing to keep in mind is, if you’re loading in your own vectors, you should base your model on en\_core\_web\_sm

In terms of using en\_core\_web\_sm, i did not see a need while saving to disk, is that ok ?

---

<div class="post-metadata">

**Author:** ![beckerfuffle](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/beckerfuffle/32/159_2.png) [@beckerfuffle](https://support.prodi.gy/u/beckerfuffle)\
**Post date:** [February 28, 2018, 4:36pm UTC](https://support.prodi.gy/t/biomedical-nlp-models-in-spacy/319/5 "2018-02-28T16:36:14Z")

</div>

Have a look here: [Loading gensim word2vec vectors for terms.teach?](https://support.prodi.gy/t/loading-gensim-word2vec-vectors-for-terms-teach/333/3)

Same use case.
