# Text does not exist in spans after NER labelling

**URL:** <https://support.prodi.gy/t/text-does-not-exist-in-spans-after-ner-labelling/6337>\
**Category:** Uncategorized\
**Tags:** ner\
**Created:** [February 2, 2023, 11:15am UTC](https://support.prodi.gy/t/text-does-not-exist-in-spans-after-ner-labelling/6337 "2023-02-02T11:15:38Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [February 2, 2023, 2:06pm UTC](https://support.prodi.gy/t/text-does-not-exist-in-spans-after-ner-labelling/6337/2 "2023-02-02T14:06:18Z")

</div>

hi @bev.manz!

> [@bev.manz](#):
>
> For some reason some of my labelled data is missing 'text' within spans entirely, I saw a similar post but they had the item 'text' but an empty string, where in my case it seems 'text' does not exist at all.
> 
> This seems to be entirely random and basically only has occured for a few entities, only have noticed after evaluating my trained model that the entities with the lowest F1 score seem to have no 'text' within the 'span' and I presume that the model would rely on 'text' also existing alongside the char span?

By default, annotated spans do not include the raw text by design.

You mention that it seems "random" that sometimes you do see the `spans` text and other times you don't. Could you be simply seeing the text for the `tokens`, not the `spans` text?

Let me describe. For example, you can view the [`ner_manual` interface](https://prodi.gy/docs/api-interfaces#ner_manual) and see what the intended annotated spans should look like:

 ![prodi.gy_docs_api-interfaces (2)](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/a/ad700ab189a8a405bbb675e6b5c05ac4d14a8358.png)

Produces:

```python
{
  "text": "First look at the new MacBook Pro",
  "spans": [
    {"start": 22, "end": 33, "label": "PRODUCT", "token_start": 5, "token_end": 6}
  ],
  "tokens": [
    {"text": "First", "start": 0, "end": 5, "id": 0},
    {"text": "look", "start": 6, "end": 10, "id": 1},
    {"text": "at", "start": 11, "end": 13, "id": 2},
    {"text": "the", "start": 14, "end": 17, "id": 3},
    {"text": "new", "start": 18, "end": 21, "id": 4},
    {"text": "MacBook", "start": 22, "end": 29, "id": 5},
    {"text": "Pro", "start": 30, "end": 33, "id": 6}
  ]
}

```

Notice that the `tokens` have the text, but not the `spans`, which are the actual annotations.

Please confirm that this is consistent with what you're seeing.

The reason is that for training, spaCy only needs the `start` and `end` info, not the actual text itself.

If you do need to add the text, you can add it with something like this:

> [@empty spans and spans with no 'text' attribute](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/10):
>
> Ah, sorry! Completely forgot. You get this if you have a record that can't find a spans. Change to this (I've also updated the code above): for eg in examples: if eg.get('spans') is not None: for span in eg.get('spans'): span['text'] = eg['text'][span['start']:span['end']] By using eg.get('spans') instead of eg["spans"] you won't get an error when it doesn't find a key. Crossing fingers that this should work crossed_fingers

---

_[View the full topic](https://support.prodi.gy/t/text-does-not-exist-in-spans-after-ner-labelling/6337)._
