# Loading pre-annotated data that has multiple sub-labels per word

**URL:** <https://support.prodi.gy/t/loading-pre-annotated-data-that-has-multiple-sub-labels-per-word/4362>\
**Category:** Uncategorized\
**Tags:** usage, spancat\
**Created:** [June 25, 2021, 1:27pm UTC](https://support.prodi.gy/t/loading-pre-annotated-data-that-has-multiple-sub-labels-per-word/4362 "2021-06-25T13:27:23Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jhodson](https://avatars.discourse-cdn.com/v4/letter/j/2acd7d/32.png) [@Jhodson](https://support.prodi.gy/u/Jhodson)\
**Post date:** [June 25, 2021, 1:27pm UTC](https://support.prodi.gy/t/loading-pre-annotated-data-that-has-multiple-sub-labels-per-word/4362/1 "2021-06-25T13:27:23Z")

</div>

Hello,

I currently have pre-annotated data that has words that require multiple labels in a hierarchical format. EX:

Text: "I took tylenol."

Tylenol - Label: Medication  
Tylenol - Sub-label: Polar  
Tylenol - Sub-label: Generic  
etc..

Currently the format to load this in a single label is:

```python
{
'text': 'I took tylenol.',
'tokens': etc.. ,
'spans':[{'start':7,'end':13,'token_start':2,'token_end':2,'label':'Medication'}]
}

```

This format loaded in using prodigy mark as a JSONL will highlight Tylenol as the medication which is a great first step. How can I edit this format to include the multiple sub-labels on the same word?

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [June 27, 2021, 2:53pm UTC](https://support.prodi.gy/t/loading-pre-annotated-data-that-has-multiple-sub-labels-per-word/4362/2 "2021-06-27T14:53:05Z")

</div>

Hi!

Traditionally, NER annotation in Prodigy allows only one label per token.

However, for Prodigy 1.11, we've created a new recipe `spans.manual` that will allow you to annotate overlapping and nested spans. Your input would look something like this (added newlines for readability but those wouldn't be in your JSONL file):

```python
{"text":"I took tylenol.",

"tokens":[{"text":"I","start":0,"end":1,"id":0,"ws":true},
{"text":"took","start":2,"end":6,"id":1,"ws":true},
{"text":"tylenol","start":7,"end":14,"id":2,"ws":false},
{"text":".","start":14,"end":15,"id":3,"ws":false}],

"spans":[{"start":7,"end":14,"token_start":2,"token_end":2,"label":"Medication"},
{"start":7,"end":14,"token_start":2,"token_end":2,"label":"Generic"}]}

```

And then with

```python
prodigy spans.manual my_output blank:en input.jsonl -l Medication,Generic

```

those spans would be preannotated:

![afbeelding](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/f/fd6e0121c20ef360df1eb758c102450c90bff049.png)

For more information on the upcoming 1.11 release, currently available as a "nightly" release, see this thread: [✨ Prodigy nightly: spaCy v3 support, UI for overlapping spans, improved feeds & more](https://support.prodi.gy/t/prodigy-nightly-spacy-v3-support-ui-for-overlapping-spans-improved-feeds-more/3861)
