# need help in creating own jsonl file for training the model

**URL:** <https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167>\
**Category:** Uncategorized\
**Tags:** usage, solved\
**Created:** [January 29, 2019, 2:00pm UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167 "2019-01-29T14:00:26Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![joshiag](https://avatars.discourse-cdn.com/v4/letter/j/45deac/32.png) [@joshiag](https://support.prodi.gy/u/joshiag)\
**Post date:** [January 29, 2019, 2:00pm UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/1 "2019-01-29T14:00:26Z")

</div>

Hello, I would like to know the structure and format for josnl file from which I can train my model. I tried formats on the website, but are not working. Its like I would like to train my model with all names of the countries, states/provinces, cities etc. Please let me know how I can achieve this?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 29, 2019, 4:36pm UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/2 "2019-01-29T16:36:36Z")

</div>

Your `PRODIGY_README.html` (available for download with Prodigy) has an “Input formats” section that shows the expected format of the data files you can load in, and a section “Annotation task formats”, which shows the format of the labelled examples Prodigy stores in the database.

Data you load in (to annotate it with a recipe like `ner.teach`) should ideally be a JSONL file with one dictionary/object per line and a `"text"` key. For example:

```json
{"text": "This is a text"}
{"text": "This is another text"}

```

What exactly are you trying to do? Which recipe are you running, and what errors did you see when you loaded your data?

---

<div class="post-metadata">

**Author:** ![joshiag](https://avatars.discourse-cdn.com/v4/letter/j/45deac/32.png) [@joshiag](https://support.prodi.gy/u/joshiag)\
**Post date:** [January 30, 2019, 6:52am UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/3 "2019-01-30T06:52:50Z")

</div>

I am trying to use {“I am from Sangli” ,{“entities”: [[11, 17 , “GPE”]]} this format with ner.make-gold and also ner.batch-train recipe, but i am getting error:  
ValueError: Failed to load task (invalid JSON).

{“I am from Sangli” ,{“entities”: [[11, 17 , “GPE” … m from Sangli" ,{“entities”: [[11, 17 , “GPE”]]}

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 30, 2019, 12:05pm UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/4 "2019-01-30T12:05:48Z")

</div>

Yes, threre are 2 problems here:

1. It’s invalid JSON. You’re using curly braces around values separated by a comma, e.g `{"foo", "bar"}`.
2. It’s not the format expected by Prodigy. Prodigy notes highlighted spans as a list of `"spans"`. See the “Annotation task formats” section in your `PRODIGY_README.html` for details. For NER, an incoming task could look like this:

```json
{
    "text": "Apple updates its analytics service with new metrics",
    "spans": [
        {"start": 0, "end": 5, "label": "ORG"}
    ]
}

```

---

<div class="post-metadata">

**Author:** ![joshiag](https://avatars.discourse-cdn.com/v4/letter/j/45deac/32.png) [@joshiag](https://support.prodi.gy/u/joshiag)\
**Post date:** [January 31, 2019, 5:58am UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/5 "2019-01-31T05:58:52Z")

</div>

Hello, Thank you for the reply. Now we tried with the format suggested by you as  
{“text”: “I am from Sangli”, “spans”: [{“start”: 11, “end”:17, “label”: “GPE”}]}  
{“text”: “I am from Satara”, “spans”: [{“start”: 11, “end”:17, “label”: “GPE”}]}  
now we are getting error with receipe  
prodigy ner.teach prodata en\_en\_pro\_web\_sm test.jsonl  
as  
File “cython\_src/prodigy/components/preprocess.pyx”, line 143, in prodigy.components.preprocess.\_add\_tokens  
KeyError: 11

Please suggest how do we proceed?

Thank you

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 31, 2019, 10:32am UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/6 "2019-01-31T10:32:47Z")

</div>

Can you try setting `--unsegmented` and see if that solves it?

---

<div class="post-metadata">

**Author:** ![joshiag](https://avatars.discourse-cdn.com/v4/letter/j/45deac/32.png) [@joshiag](https://support.prodi.gy/u/joshiag)\
**Post date:** [February 1, 2019, 6:48am UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/7 "2019-02-01T06:48:30Z")

</div>

Hello,  
Tried with --unsegmented. Now getting following error

in prodigy.models.ner.EntityRecognizer. **call**.get\_tasks.sort\_by\_entity  
KeyError: ‘start’

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 1, 2019, 9:48am UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/8 "2019-02-01T09:48:50Z")

</div>

Hmmm, I haven’t seen that erro before. Can you double-check that all entries in `"spans"` define a `"start"`, `"end"` and `"label"`?

I also wonder if what you’re trying to do will even work using `ner.make-gold` – that recipe sets its own entity annotations, so it might actually just overwrite what you already have in the data. You probably want to use `ner.manual` instead if you want to feed in pre-labelled examples.

---

<div class="post-metadata">

**Author:** ![joshiag](https://avatars.discourse-cdn.com/v4/letter/j/45deac/32.png) [@joshiag](https://support.prodi.gy/u/joshiag)\
**Post date:** [February 2, 2019, 5:12am UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/9 "2019-02-02T05:12:51Z")

</div>

As you suggested, I checked all entries in span, start, end and label. But facing same error even with make-gold by passing the jsonl to recipe. My problem is I have to train my model with very large data and manual or correcting in make-gold will consume our lot of time. So if you can help us training model with pre defined labels will save humongous efforts and the accuracy of model will be high.

thanks

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 2, 2019, 3:38pm UTC](https://support.prodi.gy/t/need-help-in-creating-own-jsonl-file-for-training-the-model/1167/10 "2019-02-02T15:38:05Z")

</div>

Oh okay – I mean, Prodigy is an annotation tool, so the point of it is always to… actually annotate data, even if it’s just double-checking. If you already have annotations and just want to train a model, you might be better off using spaCy directly. See the [training docs](https://spacy.io/usage/training) and [`spacy train`](https://spacy.io/api/cli#train) for details.
