# Convert annotated NER data to entity "offset format"

**URL:** <https://support.prodi.gy/t/convert-annotated-ner-data-to-entity-offset-format/3322>\
**Category:** Uncategorized\
**Tags:** ner, spacy, solved\
**Created:** [August 24, 2020, 4:25pm UTC](https://support.prodi.gy/t/convert-annotated-ner-data-to-entity-offset-format/3322 "2020-08-24T16:25:11Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![LBoss](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/lboss/32/1987_2.png) [@LBoss](https://support.prodi.gy/u/LBoss)\
**Post date:** [August 24, 2020, 4:25pm UTC](https://support.prodi.gy/t/convert-annotated-ner-data-to-entity-offset-format/3322/1 "2020-08-24T16:25:11Z")

</div>

Dear prodigy team,

I annotated data for NER and I want to follow the example for training from the spaCy website which can be found here:

Guides -\> Training models -\> NER -\> Updating the Named Entity Recognizer

The required input format for the trainset is:

TRAIN\_DATA = [  
("Who is Shaka Khan?", {"entities": [(7, 17, "PERSON")]}),  
("I like London and Berlin.", {"entities": [(7, 13, "LOC"), (18, 24, "LOC")]}),  
]

which is used later in nlp.update.

My question is how can I get the above described format which is required for the example ("offset format")? I used the data-to-spacy recipe but it seems to me that that format creates something else which looks like it can be used for the commandline training.

Thanks for your help!

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [August 24, 2020, 6:06pm UTC](https://support.prodi.gy/t/convert-annotated-ner-data-to-entity-offset-format/3322/2 "2020-08-24T18:06:02Z")

</div>

Hi! I think the solution might be easier than you think 🙂 When you export your annotations with `db-out`, each annotated example will contain a list of `"spans"`, and each span has a `"start"` and `"end"`. Those are the entity offsets. (You can see an example of the JSON format [here](https://prodi.gy/docs/api-interfaces#ner_manual).)

The `data-to-spacy` command produces JSON-formatted training data in spaCy's format that you can use with the `spacy train` CLI command.

---

<div class="post-metadata">

**Author:** ![LBoss](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/lboss/32/1987_2.png) [@LBoss](https://support.prodi.gy/u/LBoss)\
**Post date:** [August 25, 2020, 9:01am UTC](https://support.prodi.gy/t/convert-annotated-ner-data-to-entity-offset-format/3322/3 "2020-08-25T09:01:34Z")

</div>

> [@ines](#):
>
> with `db-out` , each annotated example will contain a list of `"spans"` , and each span has a `"start"` and `"end"`

Thanks, I extracted them from the jsonl and got the example I wanted to try running.
