# data-to-spacy --base-model usage

**URL:** <https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774>\
**Category:** Uncategorized\
**Created:** [September 7, 2023, 11:05am UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774 "2023-09-07T11:05:09Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![TatyanaKavalenkaTR](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tatyanakavalenkatr/32/3893_2.png) [@TatyanaKavalenkaTR](https://support.prodi.gy/u/TatyanaKavalenkaTR)\
**Post date:** [September 7, 2023, 11:05am UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/1 "2023-09-07T11:05:09Z")

</div>

I have the following problem ( assuming, after prodigy update).  
We uses custom tokenizer, and uses data-to-spacy utility. To be able to use our custom tokenizer, we have and uses --base-model parameter where we pass empty model (no components in pipe added) with only our custom tokenizer.

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/8/853178649513da3e956c5f13c77b82a573d4095b.png)

Not sure why we expect base model should contain tok2vec component.  
Maybe any advises how to fix ?

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [September 7, 2023, 12:11pm UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/2 "2023-09-07T12:11:31Z")

</div>

hi @TatyanaKavalenkaTR,

Thanks for your question.

Can you provide your Prodigy version? Ideally, if you can provide `prodigy stats`.

I think you may be running into this issue:

> [@Base model without tok2vec throws error](https://support.prodi.gy/t/base-model-without-tok2vec-throws-error/6576):
>
> Hello ! I have an issue since the version [v1.11.12](https://prodi.gy/docs/changelog#v1.11.12). As is stated in the doc, some bug was fixed around the --base-model usage. When I try to use a base model for NER on a simple dataset (I'm using fr\_dep\_news\_trf) prodigy returns the following error : KeyError: "[E001] No component 'tok2vec' found in pipeline. Available names: ['transformer', 'morphologizer', 'parser', 'attribute\_ruler', 'lemmatizer']" Which is... normal actually, since the fr\_dep\_news\_trf model does not have any tok2vec comp…

I think this was `transformers` related, but I'll need to check with colleagues.

Also - just curious, can you run [`spacy debug config`](https://spacy.io/api/cli#debug-config)? This helps us rule out if there's just an error with your config. Even better, if you could provide the config too, but if not, at least running debug config.

---

<div class="post-metadata">

**Author:** ![TatyanaKavalenkaTR](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tatyanakavalenkatr/32/3893_2.png) [@TatyanaKavalenkaTR](https://support.prodi.gy/u/TatyanaKavalenkaTR)\
**Post date:** [September 7, 2023, 4:23pm UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/3 "2023-09-07T16:23:51Z")

</div>

Prodigy version is 1.13.1  
Prodigy stats:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/c/c5558526286c74e8bf389731f385f547af631b3b.png)

Mentioned issue seems for me similar. But I'd like to note, my base model does not contain neither tranformers nor tok2vec

This is output for spacy debug config

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/4/4408cf7f24d2c765714e77ea3ae4ac5d92a2e86d.png)

---

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [September 8, 2023, 9:25am UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/4 "2023-09-08T09:25:56Z")

</div>

Hey @TatyanaKavalenkaTR ,

I think as a first step we should confirm what components does your `prodigy_base_model` pipeline contain. You say you "expect" it to have `tok2vec`, but we can confirm it really quick by loading it with spaCy:

```python
import spacy
nlp = spacy.load("prodigy_base_model")
nlp.pipe_names

```

which should print the names of components.

It should contain a `tok2vec` embedding layer (or a transformer embedding layer if you use transformers but that is not the case).  
If it does contain a `tok2vec` component, we need to look for a problem inside `data-to-spacy`. If it doesn't, we should look at how the `prodigy_base_model` was created. If that's the case it would be good if you could share the steps and the config of the model.  
Thanks!

---

<div class="post-metadata">

**Author:** ![TatyanaKavalenkaTR](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tatyanakavalenkatr/32/3893_2.png) [@TatyanaKavalenkaTR](https://support.prodi.gy/u/TatyanaKavalenkaTR)\
**Post date:** [September 8, 2023, 10:20am UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/5 "2023-09-08T10:20:15Z")

</div>

No, I use empty model (no pipe components added ) with only my custom tokenizer.  
So pipe\_names is an empty list.

My initial goal is to use custom tokenizer during preparing spacy files exported from prodigy dataset as it is mentioned in option description.

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/b/be2e637994f607e27c6418b2835cda3aa666b8b0.png)

Config file attached:  
[config.html](https://support.prodi.gy/uploads/short-url/3R6CQziTtGB5FXkdzRvSEMnhqRd.html) (1.9 KB)

---

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [September 13, 2023, 6:18am UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/6 "2023-09-13T06:18:53Z")

</div>

Hi @TatyanaKavalenkaTR,

Sorry for the delay in response. It's true that according to the the docs the `--base-model` is only used for tokenization and sentence segmentation but the `data-to-spacy` recipe expects the components for which the training data is being generated to be present. In your case that would be `tok2vec` and `textcat-multilabel` .  
Additionally the sourcing of the custom tokenizer is currently not automated, you'd have to provide the instruction to source it in the config file.  
Here you can find a dedicted post with examples: [Train recipe uses different Tokenizer than in ner.manual - #2 by magdaaniol](https://support.prodi.gy/t/train-recipe-uses-different-tokenizer-than-in-ner-manual/6722/2)  
It's in the context of `train` recipe but the handling of the `--base-model`parameter is the same.  
Finally, out of curiosity why do you need a custom tokenizer here?

---

<div class="post-metadata">

**Author:** ![TatyanaKavalenkaTR](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tatyanakavalenkatr/32/3893_2.png) [@TatyanaKavalenkaTR](https://support.prodi.gy/u/TatyanaKavalenkaTR)\
**Post date:** [September 13, 2023, 7:43pm UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/7 "2023-09-13T19:43:02Z")

</div>

For training I uses config from this spacy project (no tok2vec component present here) [https://github.com/explosion/projects/blob/v3/pipelines/textcat\_multilabel\_demo/configs/config.cfg](https://github.com/explosion/projects/blob/v3/pipelines/textcat_multilabel_demo/configs/config.cfg)

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/a/aef07ab0498b574441facb0395d506e120bc5742.png)

Assume, since I need to train and use only one component, there is no need in a separate tok2vec.
