# data-to-spacy --base-model usage

**URL:** <https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774>\
**Category:** Uncategorized\
**Created:** [September 7, 2023, 11:05am UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774 "2023-09-07T11:05:09Z")\
**Posts on this page:** 1\
**Showing post:** 6

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [September 13, 2023, 6:18am UTC](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774/6 "2023-09-13T06:18:53Z")

</div>

Hi @TatyanaKavalenkaTR,

Sorry for the delay in response. It's true that according to the the docs the `--base-model` is only used for tokenization and sentence segmentation but the `data-to-spacy` recipe expects the components for which the training data is being generated to be present. In your case that would be `tok2vec` and `textcat-multilabel` .  
Additionally the sourcing of the custom tokenizer is currently not automated, you'd have to provide the instruction to source it in the config file.  
Here you can find a dedicted post with examples: [Train recipe uses different Tokenizer than in ner.manual - #2 by magdaaniol](https://support.prodi.gy/t/train-recipe-uses-different-tokenizer-than-in-ner-manual/6722/2)  
It's in the context of `train` recipe but the handling of the `--base-model`parameter is the same.  
Finally, out of curiosity why do you need a custom tokenizer here?

---

_[View the full topic](https://support.prodi.gy/t/data-to-spacy-base-model-usage/6774)._
