# Implementing ner.correct says the model you are using isn't setting sentence boundaries

**URL:** <https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664>\
**Category:** Uncategorized\
**Tags:** ner, solved\
**Created:** [July 10, 2023, 9:29am UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664 "2023-07-10T09:29:47Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![tahia](https://avatars.discourse-cdn.com/v4/letter/t/b4bc9f/32.png) [@tahia](https://support.prodi.gy/u/tahia)\
**Post date:** [July 10, 2023, 9:29am UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/1 "2023-07-10T09:29:48Z")

</div>

Hi there,

I'm facing some issues in using ner.correct. My code is:

prodigy ner.correct defect\_data\_correct ./tmp\_model3/model-best sample.jsonl --label EQUIPMENT --exclude ner\_defect\_labels

ner\_defect\_labels is my manually annotated dataset which I used to train a model and then save it to tmp\_model3. The error is as follows:

Using 1 label(s): EQUIPMENT  
⚠ The model you're using isn't setting sentence boundaries (e.g. via  
the parser or sentencizer). This means that incoming examples won't be split  
into sentences.

When I open [http://localhost:8080](http://localhost:8080) to correct the annotations, it says there's nothing to annotate.

Could you please help me understand what's going wrong?

Sincerely,  
Tahia

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [July 10, 2023, 9:46am UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/2 "2023-07-10T09:46:14Z")

</div>

When you look at the [ner.correct](https://prodi.gy/docs/recipes/#ner-correct) recipe docs, you'll notice that there is a `--unsegmented` flag.

> **[Built-in Recipes · Prodigy · An annotation tool for AI, Machine Learning...](https://prodi.gy/docs/recipes/#ner-correct)**
>
> A downloadable annotation tool for NLP and computer vision tasks such as named entity recognition, text classification, object detection, image segmentation, A/B evaluation and more.

By default this recipe will split sentence on your behalf because it makes it easier to annotate examples. However, you will need to pass it a `nlp` model that is capable of doing that.

Just to check, how did you construct your model? The `./tmp_model3/model-best` one? If it was trained using `prodigy train` you may want to make sure that you use `en_core_web_sm` as a starting point. That way, you should have all the components required to split sentences.

Alternatively, you may also:

- Choose to run with the `--unsegmented` flag. This way there is no need for components that can split sentences.
- Choose to split the data into sentences beforehand using `en_core_web_sm`. This can be done in a separate Python script and you can store the sentences into a new file called `sentences.jsonl` which you then pass to `ner.correct`.

---

<div class="post-metadata">

**Author:** ![tahia](https://avatars.discourse-cdn.com/v4/letter/t/b4bc9f/32.png) [@tahia](https://support.prodi.gy/u/tahia)\
**Post date:** [July 10, 2023, 9:56am UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/3 "2023-07-10T09:56:48Z")

</div>

Thank you so much for your reply! This is how I trained the model:

prodigy train ./tmp\_model3 --ner ner\_defect\_labels --eval-split 0.3

I didn't use en\_core\_web\_sm. Would the code look like this if I did?

prodigy train ./tmp\_model3 --ner ner\_defect\_labels --eval-split 0.3 --base-model en\_core\_web\_sm

Sincerely,  
Tahia

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [July 10, 2023, 10:04am UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/4 "2023-07-10T10:04:18Z")

</div>

> [@tahia](#):
>
> prodigy train ./tmp\_model3 --ner ner\_defect\_labels --eval-split 0.3 --base-model en\_core\_web\_sm

Yep. You want to make sure this model is around beforehand though, but you should be able to simply download it via:

```python
python -m spacy download en_core_web_sm

```

---

<div class="post-metadata">

**Author:** ![tahia](https://avatars.discourse-cdn.com/v4/letter/t/b4bc9f/32.png) [@tahia](https://support.prodi.gy/u/tahia)\
**Post date:** [July 10, 2023, 1:50pm UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/5 "2023-07-10T13:50:05Z")

</div>

Thank you! ner.correct works now

---

<div class="post-metadata">

**Author:** ![tahia](https://avatars.discourse-cdn.com/v4/letter/t/b4bc9f/32.png) [@tahia](https://support.prodi.gy/u/tahia)\
**Post date:** [July 17, 2023, 11:21am UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/6 "2023-07-17T11:21:10Z")

</div>

Hi @koaning ,

I implemented ner.correct and retrained my model. The accuracy improved. However, while I was using ner.correct I noticed that my base model en\_core\_web\_sm incorrectly split the sentences. For example, a sentence that should have been a single input for annotation was split 3 times and came up in parts when I used ner.correct.

I am worried that this has incorrectly affected the model. Is there a modification I need to do to fix this?

Instead of using the base model, is it generally better to use the unsegmented tag instead?

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [July 17, 2023, 4:00pm UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/7 "2023-07-17T16:00:05Z")

</div>

You could also try other, maybe more performant, base models.

Do you have the same issues with `en_core_web_md` and `en_core_web_lg`?

---

<div class="post-metadata">

**Author:** ![tahia](https://avatars.discourse-cdn.com/v4/letter/t/b4bc9f/32.png) [@tahia](https://support.prodi.gy/u/tahia)\
**Post date:** [July 21, 2023, 1:03pm UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/8 "2023-07-21T13:03:16Z")

</div>

Hi Koaning,

I have the same issue with en\_core\_web\_lg. These are the commands I have implemented so far:

```python
prodigy ner.manual ner_defect_labels en_core_web_lg sample.jsonl --label EQUIPMENT --patterns patterns.jsonl

```

```python
db-out ner_defect_labels > annotations.jsonl

```

```python
prodigy train ./first_train --ner ner_defect_labels --eval-split 0.3 --base-model en_core_web_lg --training.max_steps=3000 --training.optimizer.learn_rate=0.001

```

```python
prodigy ner.correct first_train_correct ./first_train/model-best annotations.jsonl --label EQUIPMENT --exclude ner_defect_labels

```

In order to use ner.correct with en\_core\_web\_md, do I have to repeat from ner.manual with en\_core\_web\_md as the base model?

I'd be grateful for your help in understanding this. Please feel free to let me know if I should make a separate thread with my question.

Sincerely,  
Tahia

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [July 24, 2023, 8:32am UTC](https://support.prodi.gy/t/implementing-ner-correct-says-the-model-you-are-using-isnt-setting-sentence-boundaries/6664/9 "2023-07-24T08:32:22Z")

</div>

> [@tahia](#):
>
> In order to use ner.correct with en\_core\_web\_md, do I have to repeat from ner.manual with en\_core\_web\_md as the base model?

That's shouldn't be needed. The model that you use while annotating in `ner.manual` provides the tokenisation and the sentence-splitting capabilities. The tokenisers are the same across all `en_core_*` models.

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/f/fbdc8881221e3ff913a3ab672ffd03834730c6d2.png)

If you're eager to learn more, you may find [this section useful from the spaCy documentation](https://spacy.io/usage/processing-pipelines).

The only slight difference is that the sentences may be split somewhat differently. I'd be surprised if that has a significant negative impact on your final model though, mainly because you're still training on examples that you've accepted.

Does this help?
