# Ner format to CONLL

**URL:** <https://support.prodi.gy/t/ner-format-to-conll/1153>\
**Category:** Uncategorized\
**Tags:** usage, ner, solved\
**Created:** [January 24, 2019, 2:37pm UTC](https://support.prodi.gy/t/ner-format-to-conll/1153 "2019-01-24T14:37:02Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![JoaoMVR](https://avatars.discourse-cdn.com/v4/letter/j/ba8739/32.png) [@JoaoMVR](https://support.prodi.gy/u/JoaoMVR)\
**Post date:** [January 24, 2019, 2:37pm UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/1 "2019-01-24T14:37:02Z")

</div>

I’ve just acquired prodigy to work on a manual ner tagging task. After manually tagging a document, the tool exports the results to a json format, the stantard spacy format. We would like to have this data tagged in the CoNLL format, the column format like so:

John PERSON  
works O  
for O  
Microsoft ORGANIZATION

Is there an option to do this or should we opt to post-process the json file in order to do so?

Thanks in advance

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 24, 2019, 4:07pm UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/2 "2019-01-24T16:07:44Z")

</div>

Hi! I’d recommend writing your own converter, yes. spaCy actually ships with a [`biluo_tags_from_offsets`](https://spacy.io/api/goldparse#biluo_tags_from_offsets) helper that takes a text and character offsets and returns the BILUO entity labels. So this might be helpful?

You can also interact with Prodigy’s database directly from Python, so you’ll be able to skip the whole exporting/importing/exporting part.

Here’s an example (untested, but something along those lines should work):

```python
from prodigy.components.db import connect
from spacy.gold import biluo_tags_from_offsets
from spacy.lang.en import English # or whichever language tokenizer you need

nlp = English()

db = connect() # uses settings from your prodigy.json
examples = db.get_dataset('your_dataset') # load the annotations

for eg in examples:
    doc = nlp(eg['text'])
    entities = [(span['start'], span['end'], span['label'])
                for span in eg['spans']]
    tags = biluo_tags_from_offsets(doc, entities)
    # do something with the tags here

```

---

<div class="post-metadata">

**Author:** ![JoaoMVR](https://avatars.discourse-cdn.com/v4/letter/j/ba8739/32.png) [@JoaoMVR](https://support.prodi.gy/u/JoaoMVR)\
**Post date:** [January 24, 2019, 4:32pm UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/3 "2019-01-24T16:32:15Z")

</div>

Works perfectly, thanks a lot!

---

<div class="post-metadata">

**Author:** ![bjornvandijkman](https://avatars.discourse-cdn.com/v4/letter/b/ed655f/32.png) [@bjornvandijkman](https://support.prodi.gy/u/bjornvandijkman)\
**Post date:** [June 4, 2019, 9:43am UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/4 "2019-06-04T09:43:07Z")

</div>

I have tried to do this for my jsonl file. However, I’m quite new to programming. The following script almost works, but it seems to go over the same line in the json file 8 times. Can anyone point out to me what I’m doing wrong here? `Result` is my data here.

```
extended_tags = []
extended_entities = []
extended_token = []

for i in range(len(result)):
    data = result[i]
    for d in data:
        doc = nlp(data['text'])
        for token in doc:
            extended_token.append(token)
        entities = [(span['start'], span['end'], span['label'])
        for span in data['spans']]
        tags = biluo_tags_from_offsets(doc, entities)
        extended_tags.extend(tags)             
        extended_entities.extend(entities)
```

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 4, 2019, 9:52am UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/5 "2019-06-04T09:52:16Z")

</div>

@bjornvandijkman I think the indentation got a bit messed up when you copy-pasted the code over. Could you update that when you have a second? Otherwise, it’s a bit difficult to follow because it’s unclear what’s in which block. Also, what does your `result` look like?

---

<div class="post-metadata">

**Author:** ![bjornvandijkman](https://avatars.discourse-cdn.com/v4/letter/b/ed655f/32.png) [@bjornvandijkman](https://support.prodi.gy/u/bjornvandijkman)\
**Post date:** [June 4, 2019, 10:06am UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/6 "2019-06-04T10:06:22Z")

</div>

I updated the code. The result is a dataset created using `ner.manual` and imported to `examples` as you indicated. Then the only thing I did was removed the examples where I did not `accept` the annotation using the following code:

```
# Only keep the accepted answers, as the rejected ones have no span
result = []
for i in examples:
    if i['answer'] == "accept":
        result.append(i)

```

Format looks as follows for two of the lines:

> text  
> \_input\_hash  
> \_task\_hash  
> tokens  
> \_session\_id  
> \_view\_id  
> spans  
> answer  
> text  
> \_input\_hash  
> \_task\_hash  
> tokens  
> \_session\_id  
> \_view\_id  
> spans  
> answer

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 4, 2019, 10:16am UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/7 "2019-06-04T10:16:32Z")

</div>

Thanks! I think the problem is this: `for d in data:`. In the outer loop, you’re going over each example in your dataset, which is a dictionary and which you’re storing in the variable `data` in your code.

However, you _then_ go and _also_ iterate over `data` and parse the text and create the tags each time. So basically, instead of doing it once per example, you’re doing it once for each key in your example dict. Iterating over a dict in Python is perfectly valid and what happens is that you iterating over the keys. For instance:

```python
data = {"text": "hello", "meta": "world"}
for d in data:
    print(d)

```

This will print `text` and `meta`. Your example dict happens to have 8 keys, so your loop runs 8 times per example.

**TL;DR:** Remove the `for d in data:`, you don’t need that.

Btw, another tip: If you have a list (like your examples in `result`), you can also just iterate over its elements in the `for` loop, instead of indexing into it. For example, instead of this:

```python
for i in range(len(result)):
    data = result[i]

```

… you can write this:

```python
for data in result:

```

---

<div class="post-metadata">

**Author:** ![bjornvandijkman](https://avatars.discourse-cdn.com/v4/letter/b/ed655f/32.png) [@bjornvandijkman](https://support.prodi.gy/u/bjornvandijkman)\
**Post date:** [June 4, 2019, 10:24am UTC](https://support.prodi.gy/t/ner-format-to-conll/1153/8 "2019-06-04T10:24:01Z")

</div>

Thank you for being so patient with me and thanks for the help! The support section is truly amazing here 🙂
