# Train recipe uses different Tokenizer than in ner.manual

**URL:** <https://support.prodi.gy/t/train-recipe-uses-different-tokenizer-than-in-ner-manual/6722>\
**Category:** Uncategorized\
**Tags:** ner\
**Created:** [August 8, 2023, 7:23am UTC](https://support.prodi.gy/t/train-recipe-uses-different-tokenizer-than-in-ner-manual/6722 "2023-08-08T07:23:41Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [August 8, 2023, 1:15pm UTC](https://support.prodi.gy/t/train-recipe-uses-different-tokenizer-than-in-ner-manual/6722/2 "2023-08-08T13:15:43Z")

</div>

Welcome to the forum @kaiser 👋

Sourcing the custom tokenizer when passing `--base_model` is currently not automated. You'd need to modify it via the `config.cfg` file.  
If you already have your modified base\_model as package, you could try adding the following to the config file:

```python
# Inside your .cfg file
...
[initialize.before_init]
@callbacks = "spacy.copy_from_base_model.v1"
tokenizer = "your_base_model"
vocab = "your_base_model"
...

```

Alternatively, you can:

1. Define your modification as a registered callback:

```python
# functions.py
from spacy.util import registry, compile_infix_regex
from spacy.lang.char_classes import ALPHA, ALPHA_LOWER, ALPHA_UPPER, HYPHENS

@registry.callbacks("customize_tokenizer")
def make_customize_tokenizer():
    def customize_tokenizer(nlp):
        infix_rules = nlp.Defaults.infixes + [r"(?<=[{a}])(?:{h})(?=[{a}])".format(a=ALPHA, h=HYPHENS)]
        infix_re = spacy.util.compile_infix_regex(infix_rules)
        nlp.tokenizer.infix_finditer = infix_re.finditer
    return customize_tokenizer

```

1. Reference this callback in the config file so that it runs before the pipeline initialization

```python
[initialize]

[initialize.before_init]
@callbacks = "customize_tokenizer"

```

See spaCy docs for details: [Training Pipelines & Models · spaCy Usage Documentation](https://spacy.io/usage/training#custom-tokenizer)

Then you'd run `train`like so:  
`python -m prodigy train ./output -n ner_dataset --config config.cfg -F functions.py`  
where `functions.py` contains the registered callback for modifying the tokenizer.

Please note that you can generate the config with `spacy init config`[https://spacy.io/usage/training/#quickstart](https://spacy.io/usage/training/#quickstart)

---

_[View the full topic](https://support.prodi.gy/t/train-recipe-uses-different-tokenizer-than-in-ner-manual/6722)._
