# Anotation task format for ner\_manual interface

**URL:** <https://support.prodi.gy/t/anotation-task-format-for-ner-manual-interface/1501>\
**Category:** Uncategorized\
**Tags:** usage, ner, solved\
**Created:** [May 9, 2019, 6:50am UTC](https://support.prodi.gy/t/anotation-task-format-for-ner-manual-interface/1501 "2019-05-09T06:50:45Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [May 9, 2019, 1:41pm UTC](https://support.prodi.gy/t/anotation-task-format-for-ner-manual-interface/1501/4 "2019-05-09T13:41:51Z")

</div>

Ah, okay, that makes sense then. Are you just using `mark`? I think I misread your initial question and thought you were using the built-in `ner.manual` recipe, which does take care of the tokenization automatically.

If you need your own custom tokens that align with your entity spans, then you also need to provide them. It might be worth writing a little script to check how many of the spans do not align – maybe it’s just one or two that you can easily correct manually (or exclude from your data).

An easy way to do this is to use spaCy’s [`Doc.char_span` method](https://spacy.io/api/doc#char_span), which creates a token span from character offsets. If the character offsets don’t align to the tokens, it returns `None`. So you can do something like this:

```python
nlp = spacy.load("en_core_web_sm") # or other model

for example in examples: # your existing examples
    doc = nlp(example["text"])
    for span in example["spans"]:
        char_span = doc.char_span(span["start"], span["end"])
        if char_span is None: # start and end don't map to tokens
            print("Misaligned tokens", example["text"], span)

```

---

_[View the full topic](https://support.prodi.gy/t/anotation-task-format-for-ner-manual-interface/1501)._
