# How to modify the tokenizer used by Prodigy's recipes?

**URL:** <https://support.prodi.gy/t/how-to-modify-the-tokenizer-used-by-prodigys-recipes/445>\
**Category:** Uncategorized\
**Tags:** usage, spacy\
**Created:** [March 27, 2018, 1:30pm UTC](https://support.prodi.gy/t/how-to-modify-the-tokenizer-used-by-prodigys-recipes/445 "2018-03-27T13:30:45Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![daniel](https://avatars.discourse-cdn.com/v4/letter/d/3ec8ea/32.png) [@daniel](https://support.prodi.gy/u/daniel)\
**Post date:** [March 27, 2018, 1:30pm UTC](https://support.prodi.gy/t/how-to-modify-the-tokenizer-used-by-prodigys-recipes/445/1 "2018-03-27T13:30:45Z")

</div>

When using `ner.manual` or `ner.make_gold` recipes, sometimes the tokenizer is not giving the desired results. For example, the following text:

> How to play Pink Floyd- "Wish You Were Here"

Contains a typo "-" (which should have space before it), but the tokenizer creates the tokens `Pink` and `Floyd-`, so there is no way to mark the entity `Pink Floyd` in the Prodigy UI without the trailing "-".

I understand there is a way to customize the tokenizer's prefixes/suffixes in code but how can I do it to work with Prodigy's recipes?

Is there a way to update the model's prefixes/suffixes list and save it back to the model? So anything that uses the new model correctly tokenizes the example above without the need to create a custom tokenizer each time in code?

Thanks.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 27, 2018, 3:17pm UTC](https://support.prodi.gy/t/how-to-modify-the-tokenizer-used-by-prodigys-recipes/445/2 "2018-03-27T15:17:06Z")

</div>

> [@daniel](#):
>
> Is there a way to update the model’s prefixes/suffixes list and save it back to the model? So anything that uses the new model correctly tokenizes the example above without the need to create a custom tokenizer each time in code?

Yes, this is actually the main reason Prodigy recipes always take a model and don't just use the tokenizer or language classes shipped with spaCy directly. The idea is that you can create a custom base model for your own "dialect" of English (music language on YouTube).

When you save out a model via `nlp.to_disk`, the tokenizer is also serialized. This includes the prefix/suffix/infix rules, as well as the tokenizer exceptions. [See here](https://github.com/explosion/spaCy/blob/e0ae39060778003fcbc81a7b149962e7c096abfa/spacy/tokenizer.pyx#L352) for the respective code in `tokenizer.pyx`. So you could overwrite `nlp.tokenizer.suffix_search` with your own function and then save out the model again.

---

<div class="post-metadata">

**Author:** ![daniel](https://avatars.discourse-cdn.com/v4/letter/d/3ec8ea/32.png) [@daniel](https://support.prodi.gy/u/daniel)\
**Post date:** [March 27, 2018, 7:34pm UTC](https://support.prodi.gy/t/how-to-modify-the-tokenizer-used-by-prodigys-recipes/445/3 "2018-03-27T19:34:25Z")

</div>

> [@ines](#):
>
> When you save out a model via nlp.to\_disk, the tokenizer is also serialized. This includes the prefix/suffix/infix rules, as well as the tokenizer exceptions. See here for the respective code in tokenizer.pyx. So you could overwrite nlp.tokenizer.suffix\_search with your own function and then save out the model again.

That worked perfectly, thanks.
