# Re-labling custom dataset with Prodigy

**URL:** <https://support.prodi.gy/t/re-labling-custom-dataset-with-prodigy/4366>\
**Category:** Uncategorized\
**Tags:** usage, ner\
**Created:** [June 26, 2021, 10:43am UTC](https://support.prodi.gy/t/re-labling-custom-dataset-with-prodigy/4366 "2021-06-26T10:43:07Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![AnastKuz](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/anastkuz/32/2218_2.png) [@AnastKuz](https://support.prodi.gy/u/AnastKuz)\
**Post date:** [June 26, 2021, 10:43am UTC](https://support.prodi.gy/t/re-labling-custom-dataset-with-prodigy/4366/1 "2021-06-26T10:43:07Z")

</div>

Hello!  
I was wondering if I can load my own dataset (text and labels - IOB scheme at the moment) in Prodigy and if labeled entities will be highlighted?  
Text was labeled with prediction by PyTorch model and I wanted to load it in Prodigy, to see labels already highlighted and to check if everything is correctly labeled and if not - correct it.  
Is it possible? At the moment I have two files - one with text (sentence per line), the other one only with labels of type: (O O O B-MTRL I-MTRL O O) per line. Do I need to change format to dictionary of text, tokens, spans before loading it with db-in?

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [June 27, 2021, 3:33pm UTC](https://support.prodi.gy/t/re-labling-custom-dataset-with-prodigy/4366/2 "2021-06-27T15:33:22Z")

</div>

Hi!

I think the best option for you would be to further preprocess the already labeled texts that you have, convert them into Prodigy format and then use `ner.manual` to correct them.

The target output format you want to obtain is something like the following. I added newlines for readability here, but ideally you'd have this in a JSONL file without any enters per text example:

```python
{"text":"We are visiting London and Berlin tomorrow", 

"tokens":[{"text":"We","start":0,"end":2,"id":0},
{"text":"are","start":3,"end":6,"id":1},
{"text":"visiting","start":7,"end":15,"id":2},
{"text":"London","start":16,"end":22,"id":3},
{"text":"and","start":23,"end":26,"id":4},
{"text":"Berlin","start":27,"end":33,"id":5},
{"text":"tomorrow","start":34,"end":42,"id":6}], 

"spans": [{"start":16,"end":22,"label":"CITY"},
{"start":27,"end":33,"label":"CITY"}]}

```

Then if you'd run

```python
 prodigy ner.manual output_db blank:en input_annotated_texts.jsonl -l CITY

```

You'd get:  
 ![afbeelding](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/2/268cccdcca663d70a802297266ab067b0e2a2108.png)

And then you can either hit "accept" if it's all good, or correct the annotations/labels first.

To get to the required JSONL format from your IOB annotations, you can use `spaCy` for the conversion, e.g.:

```python

    vocab = English().vocab
    doc = Doc(vocab, words=["We", "are", "visiting", "London", "and", "Berlin", "tomorrow"], spaces=[True, True, True, True, True, True, False], ents=["O", "O", "O", "B-CITY", "O", "B-CITY", "O"])
    for ent in doc.ents:
        print(ent.start_char, ent.end_char, ent.label_)

```

Will give you

```python
16 22 CITY
27 33 CITY

```

Or have a look at some of spaCy's built-in utility tools, eg [https://spacy.io/api/top-level#biluo\_tags\_to\_spans](https://spacy.io/api/top-level#biluo_tags_to_spans)

Hope that helps you get started in the right direction! 🙂

---

<div class="post-metadata">

**Author:** ![AnastKuz](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/anastkuz/32/2218_2.png) [@AnastKuz](https://support.prodi.gy/u/AnastKuz)\
**Post date:** [June 28, 2021, 7:49am UTC](https://support.prodi.gy/t/re-labling-custom-dataset-with-prodigy/4366/3 "2021-06-28T07:49:13Z")

</div>

Thanks a lot for the detailed answer! 🙂
