# NER document Labeling

**URL:** <https://support.prodi.gy/t/ner-document-labeling/1741>\
**Category:** Uncategorized\
**Tags:** ner, solved\
**Created:** [July 4, 2019, 10:42pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741 "2019-07-04T22:42:04Z")\
**Posts on this page:** 1\
**Showing post:** 19

<div class="post-metadata">

**Author:** ![mystuff](https://avatars.discourse-cdn.com/v4/letter/m/a88e4f/32.png) [@mystuff](https://support.prodi.gy/u/mystuff)\
**Post date:** [July 15, 2019, 3:09pm UTC](https://support.prodi.gy/t/ner-document-labeling/1741/19 "2019-07-15T15:09:29Z")

</div>

I have annotated 150 htmls, the raw text is separated by new lines. Now i need to run "prodigy ner.batch-train ". am i right?.

python -m prodigy ner.batch-train company\_details\_dataset en\_core\_web\_sm --output company\_model --label COMPANY\_TYPE,COMPANY\_INFORMATION, COMPANY\_NAME,COMPANY\_DEPARTMENT,COMPANY\_ADDRESS,COMPANY\_COUNTRY\_USA,EMAIL,NAME

when i run above train command that i am getting below error:  
File "transition\_system.pyx", line 148, in spacy.syntax.transition\_system.TransitionSystem.set\_costs  
ValueError: [E024] Could not find an optimal move to supervise the parser. Usually, this means the GoldParse was not correct. For example, are all labels added to the model?

It seems that there are some white spaces issue so i followed below post to fix it.

> [@ner.batch-train after ner.maual results error (Value error : \[E024\])](https://support.prodi.gy/t/ner-batch-train-after-ner-maual-results-error-value-error-e024/1677/6):
>
> You can export your dataset by running the db-out command and then check the JSONL file: prodigy db-out resume\_ner \> resume\_ner.jsonl After you’ve removed the problematic spans or have corrected them, you can then reimport the data to a new dataset: prodigy db-in resume\_ner\_fixed resume\_ner.jsonl You can probably also write a script to find the problematic entities automatically and then exclude them, and add the result to a new dataset. I haven’t tested this yet, but something like this sh…

Still getting same error. I can see many of these in the dataset {"text":"\n","start":3359,"end":3360,"id":594},. do you think its a tokenization issue?. If so, how do i pass new line tokenizer while running "prodigy ner.batch-train "

---

_[View the full topic](https://support.prodi.gy/t/ner-document-labeling/1741)._
