# NER document Labeling

**URL:** <https://support.prodi.gy/t/ner-document-labeling/1741>\
**Category:** Uncategorized\
**Tags:** ner, solved\
**Created:** [July 4, 2019, 10:42pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741 "2019-07-04T22:42:04Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 4, 2019, 10:42pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/1 "2019-07-04T22:42:04Z")

</div>

Hi, I am new top to prodigy and i have used dataturcks a lot for labeling.  
I need to extract organization name, location, email,contact name from contact us page of a given company HTML file . I am thinking of workflow like this.  
-\>download around 50 html sources  
-\>remove noise in the html source like remove footer, input,img,script, style etc  
-\>extract remaining text data and store it in text file which contains data in separate lines.

Each cleaned html text file contains multiple sentences separated by new lines. Now i need to label each text file. I want to use ner.manual to label the data. Can someone clarify few things for me.

1. how to uniquely identify and label each document?.
2. i need to convert that text file into JSON or JSONL. Do i need to dump one cleaned file into one json file? or keep all 50 htmls data into one big json file like below?

```json
    "data": [
        {
            "text": "Apple Online Store
                        Visit the Apple Online Store to purchase Apple hardware, software and third-party accessories. To purchase by phone, please call 0800 048 0408. Lines are open Monday-Friday 08:00-20:00 and Saturday-Sunday 09:00-18:00.
.....................................
................................",
            "text": "Helpline &amp; Contact | Samsung UK
                          By ticking this box, I accept Samsung Service Updates, including : samsung.com Services and marketing information, new product and service announcements as well as special offers, events and newsletters
MOBILE: 24 HOURS, 7 DAYS A WEEK</p><p>ALL OTHER: M-F 8–12AM/S-S 9AM–11PM, APPLIANCES 6PM ET"
          
        }
    ]

```

1. can i feed data set to non-spacy api as well?

---

<div class="post-metadata">

**Author:** ![jsnleong](https://avatars.discourse-cdn.com/v4/letter/j/b9bd4f/32.png) [@jsnleong](https://support.prodi.gy/u/jsnleong)\
**Post date:** [July 5, 2019, 1:37am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/2 "2019-07-05T01:37:10Z")

</div>

Hi, this is my thought process..

> [@mystuff](#):
>
> how to uniquely identify and label each document?

I am assuming you have 50 docs in total. So you can simply write a function that takes each html text file and tagging them to a generic term such as "1000.txt", "1001.txt", 1002.txt", "1003.txt", ...  
Each sentences (separated by new lines) will still have the same doc label so long it falls under the same doc. Q2 answers the format that you require your data set to be in.

> [@mystuff](#):
>
> need to convert that text file into JSON or JSONL. Do i need to dump one cleaned file into one json file? or keep all 50 htmls data into one big json file like below?

Yes you will dump all 50 cleaned html files into one JSONL file. As mentioned above, each html file are differentiated from another by the "source" key it comes from. So, it should be in the following format...  
{"text": " XXXXXXX ", "meta": {"source" : "1000.txt"}}

Hope it helps.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 7, 2019, 6:42pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/3 "2019-07-07T18:42:03Z")

</div>

Yes, @jsnleong's solution for converting the data should work 🙂

> [@mystuff](#):
>
> I need to extract organization name, location, email,contact name from contact us page of a given company HTML file .

This sounds like you definitely want to frame this as an NER task: label spans of text in your data for the different labels, and then train a model to reproduce this decision. The most straightforward way would be to run `ner.manual` with your labels:

```bash
prodigy ner.manual your_dataset en_core_web_sm your_converted_data.jsonl --label ORG,LOCATION,EMAIL,NAME

```

Prodigy also encourages you to find more clever ways to automate the annotation so you have to do less work manually. For instance, one you have a pre-trained model that predicts _something_, you can have the model pre-highlight the entities. That's what workflows like the `ner.make-gold` are designed for.

> [@mystuff](#):
>
> can i feed data set to non-spacy api as well?

Sure! You can run the `db-out` command to export your annotations to a JSONL file, and then use that to train pretty much any model using any framework. Prodigy uses a pretty straightforward JSONL format for the created annotations that should hopefully be very easy to use an work with. Here's an example of an annotated text with an entity:

```json
{
    "text": "Hello Apple",
    "tokens": [
        { "text": "Hello", "start": 0, "end": 5, "id": 0 },
        { "text": "Apple", "start": 6, "end": 11, "id": 1 }
    ],
    "spans": [{ "start": 6, "end": 11, "label": "ORG", "token_start": 1, "token_end": 1 }]
}

```

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 8, 2019, 7:33am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/4 "2019-07-08T07:33:04Z")

</div>

Thank you so much for your prompt reply. Now i have created jsonl file with 100 json content. I am about to annotate by running the manual recipe.

python -m prodigy ner.manual company\_details\_dataset en\_core\_web\_sm your\_converted\_data.jsonl --label COMPANY\_TYPE,COMPANY\_INFORMATION, COMPANY\_NAME,COMPANY\_DEPARTMENT,COMPANY\_ADDRESS,COMPANY\_COUNTRY\_USA,EMAIL,NAME.

But i am getting following error on web ui. I think i need to add my labels to the model. May i know how to add my own lables to the model?

ERROR: can’t fetch tasks. Make sure the server is running correctly.  
Oops, something went wrong ☹

I have followed ["No tasks available" for any text source I give for ner.teach recipe](https://support.prodi.gy/t/no-tasks-available-for-any-text-source-i-give-for-ner-teach-recipe/92/3) but i can not use built-in labels for my task as i have to identify type of the company, categorize different addresses and also different sections of company name like Departments.

LOG:  
File “cython\_src\prodigy\core.pyx”, line 130, in prodigy.core.Controller.get\_questions  
File “cython\_src\prodigy\components\feeds.pyx”, line 58, in prodigy.components.feeds.SharedFeed.get\_questions  
File “cython\_src\prodigy\components\feeds.pyx”, line 63, in prodigy.components.feeds.SharedFeed.get\_next\_batch  
File “cython\_src\prodigy\components\feeds.pyx”, line 147, in prodigy.components.feeds.SessionFeed.get\_session\_stream  
ValueError: Error while validating stream: no first example. This likely means that your stream is empty.  
Task queue depth is 1  
Exception when serving /get\_session\_questions  
Traceback (most recent call last):  
File “cython\_src\prodigy\components\feeds.pyx”, line 140, in prodigy.components.feeds.SessionFeed.get\_session\_stream  
File “C:\anaconda3\lib\site-packages\toolz\itertoolz.py”, line 368, in first  
return next(iter(seq))  
StopIteration  
…

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 8, 2019, 7:52am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/6 "2019-07-08T07:52:03Z")

</div>

When you see the error “Error while validating stream: no first example. This likely means that your stream is empty.”, this usually means that there’s nothing valid to load and that the incoming stream of examples is empty. What does `your_converted_data.jsonl` look like? It should be a valid JSONL file with every record containing a `"text"`. For example:

```json
{"text": "hello world"}
{"text": "this is a text"}

```

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 8, 2019, 8:13am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/7 "2019-07-08T08:13:39Z")

</div>

{“text”: “xxxxxxxxx”, “meta”: {“source”: “company1.txt”}}  
{“text”: “yyyyyyyyy”, “meta”: {“source”: “company2.txt”}}  
{“text”: “zzzzzzzzz”, “meta”: {“source”: “company3.txt”}}  
{“text”: “aaaaaaaaa”, “meta”: {“source”: “company4.txt”}}

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 8, 2019, 8:16am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/8 "2019-07-08T08:16:28Z")

</div>

That looks correct! Could you run the command with `PRODIGY_LOGGING=basic` and share the output? So basically:

```python
PRODIGY_LOGGING=basic python -m prodigy ner.manual company_details_dataset en_core_web_sm your_converted_data.jsonl --label COMPANY_TYPE,COMPANY_INFORMATION, COMPANY_NAME,COMPANY_DEPARTMENT,COMPANY_ADDRESS,COMPANY_COUNTRY_USA,EMAIL,NAME

```

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 8, 2019, 8:20am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/9 "2019-07-08T08:20:11Z")

</div>

‘PRODIGY\_LOGGING’ is not recognized as an internal or external command,  
operable program or batch file.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 8, 2019, 8:21am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/10 "2019-07-08T08:21:30Z")

</div>

What operating system are you on? In any case, you have to set the environment variable `PRODIGY_LOGGING` to `basic` – so if you google “set environment variable” plus your OS / environment, it should tell you how to do it 🙂

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 8, 2019, 10:21am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/11 "2019-07-08T10:21:38Z")

</div>

My OS: windows  
set the environment variables below. Still cant recognize. Restarted the machine too.  
PRODIGY\_HOME=C:\Users\aaa.bbb\.prodigy  
PRODIGY\_LOGGING=basic

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 8, 2019, 10:35am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/12 "2019-07-08T10:35:34Z")

</div>

I think you might have to call `set`? See here:

> <https://superuser.com/questions/79612/setting-and-getting-windows-environment-variables-from-the-command-prompt>

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 8, 2019, 10:48am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/13 "2019-07-08T10:48:23Z")

</div>

No I meant i did set properly with set at first instance (before your message). Everything is in environment variables. Still not recognized. Sorry for bothering.

12:01:15 - GET: /project  
Task queue depth is 1  
Task queue depth is 1  
12:01:15 - POST: /get\_session\_questions  
12:01:15 - FEED: Finding next batch of questions in stream  
12:01:15 - CONTROLLER: Validating the first batch for session: data\_100-default  
12:01:15 - PREPROCESS: Tokenizing examples  
12:01:15 - FILTER: Filtering duplicates from stream  
12:01:15 - FILTER: Filtering out empty examples for key ‘text’  
Exception when serving /get\_session\_questions

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 8, 2019, 1:32pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/14 "2019-07-08T13:32:01Z")

</div>

I solved this my upgrading murmurhash.

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 8, 2019, 3:53pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/15 "2019-07-08T15:53:46Z")

</div>

![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/1X/abaf7bc1c1771193fd08c21cba40992ee7bb0241.png)

1)There are new lines symbols at the end of each line. is it common in Podigy?.

2)Also if there is a paragraph with new lines which i need to tag it as COMPANY\_INFORMATION, but there is a contact information inside the paragraph. So i need to do label3 inside label4. But UI is not allowing me to do. is it possible to configure Prodigy somewhere to allow that option?

3)\*\*555 \*\*  
**Bloemfontein**  
**South Africa**  
can i label those three lines as one label called COMPANY\_ADDRESS.?. OR does it need to be in one line?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 8, 2019, 4:26pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/16 "2019-07-08T16:26:51Z")

</div>

There are several things here: Yes, knowing where a newline is is usually very important when you're annotating named entities. Newlines are tokens and you never want to accidentally highlight them (and without the symbols, they'd be pretty much invisible). You can hide them by setting `"hide_true_newline_tokens": false` in your `prodigy.json`, but I usually wouldn't recommend it, because it can easily lead to inconsistent annotations.

Alternatively, you might also consider preprocessing that normalises the whitespace. If you're training a model later on, just make sure to also pre-process your inputs at runtime to make sure it matches the training data.

> [@mystuff](#):
>
> So i need to do label3 inside label4. But UI is not allowing me to do. is it possible to configure Prodigy somewhere to allow that option?

If you want to train a named entity recognition model (especially with spaCy), training it to predict overlapping spans isn't possible. By definition, a token can only be part of one entity. That's also why you can't highlight overlapping spans. You can always make several passes over the data to capture nested spans, but I'm not sure that's the best solution here. It really depends on what you want to do with the data later on and what statistical model you want to train.

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 10, 2019, 5:15pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/17 "2019-07-10T17:15:34Z")

</div>

Thanks for your input. I was thinking of replacing new lines with spaces. Then the content will be in one big chunk of paragraph to label. Do you think its a good idea?. As i need to label company information as well which some times multi line information which is hard to label.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 10, 2019, 5:29pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/18 "2019-07-10T17:29:07Z")

</div>

@mystuff You don’t necessarily have to replace all newlines – you’d just have to make sure that the tokenizer produces separate tokens for newlines. For example, by replacing double newlines with single newlines. Or you could add a custom tokenization rule that always splits on `\n`.

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 15, 2019, 3:09pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/19 "2019-07-15T15:09:29Z")

</div>

I have annotated 150 htmls, the raw text is separated by new lines. Now i need to run "prodigy ner.batch-train ". am i right?.

python -m prodigy ner.batch-train company\_details\_dataset en\_core\_web\_sm --output company\_model --label COMPANY\_TYPE,COMPANY\_INFORMATION, COMPANY\_NAME,COMPANY\_DEPARTMENT,COMPANY\_ADDRESS,COMPANY\_COUNTRY\_USA,EMAIL,NAME

when i run above train command that i am getting below error:  
File "transition\_system.pyx", line 148, in spacy.syntax.transition\_system.TransitionSystem.set\_costs  
ValueError: [E024] Could not find an optimal move to supervise the parser. Usually, this means the GoldParse was not correct. For example, are all labels added to the model?

It seems that there are some white spaces issue so i followed below post to fix it.

> [@ner.batch-train after ner.maual results error (Value error : \[E024\])](https://support.prodi.gy/t/ner-batch-train-after-ner-maual-results-error-value-error-e024/1677/6):
>
> You can export your dataset by running the db-out command and then check the JSONL file: prodigy db-out resume\_ner \> resume\_ner.jsonl After you’ve removed the problematic spans or have corrected them, you can then reimport the data to a new dataset: prodigy db-in resume\_ner\_fixed resume\_ner.jsonl You can probably also write a script to find the problematic entities automatically and then exclude them, and add the result to a new dataset. I haven’t tested this yet, but something like this sh…

Still getting same error. I can see many of these in the dataset {"text":"\n","start":3359,"end":3360,"id":594},. do you think its a tokenization issue?. If so, how do i pass new line tokenizer while running "prodigy ner.batch-train "

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 16, 2019, 8:20am UTC](https://support.prodi.gy/t/ner-document-labeling/1741/20 "2019-07-16T08:20:19Z")

</div>

> [@mystuff](#):
>
> Still getting same error. I can see many of these in the dataset {“text”:"\n",“start”:3359,“end”:3360,“id”:594},. do you think its a tokenization issue?. If so, how do i pass new line tokenizer while running "prodigy ner.batch-train "

Tokens containing `\n` are totally fine – it's only a problem if labelled entity spans in the `"spans"` start or end with a newline token, or consist only of newline tokens. This is an explicit change to the entity recognizer in spaCy v2.1 to make it more accurate and to prevent it from predicting entity spans like this, which are usually never what you want.

So if your data contains entries in the `"spans"` that are invalid like that, you should be able to just remove them and then re-import the edited data to a new dataset.

---

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 17, 2019, 4:14pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/21 "2019-07-17T16:14:10Z")

</div>

Thanks for your reply. I have found 3 spans with white spaces and \n and i removed and reloaded back to completley new dataset. Still getting same error.

Just for testing i tested with top4 and didnt get any error.

python -m prodigy ner.batch-train top4\_ner en\_core\_web\_sm --output top4-model --eval-split 0.2 --n-iter 6 --dropout 0.2  
The output for top4:  
17:09:14 - MODEL: Merging entity spans of 0 examples  
17:09:14 - MODEL: Using 0 examples (without ‘ignore’)  
17:09:14 - MODEL: Evaluated 0 examples  
06 1884.572 0 0 0 0 0.000

Correct 0  
Incorrect 0  
Baseline 0.000  
Accuracy 0.000

17:09:14 - RECIPE: Restoring disabled pipes: [‘tagger’, ‘parser’]

Model: C:\top4-model  
Training data: C:\top4-model\training.jsonl

does it look alright for you?. if so, can i redo entire annotation for those 100 .

[Next page](https://support.prodi.gy/t/ner-document-labeling/1741.md?page=2)
