# Prodigy Train NER Results explanation

**URL:** <https://support.prodi.gy/t/prodigy-train-ner-results-explanation/4400>\
**Category:** Uncategorized\
**Tags:** usage, ner, solved\
**Created:** [July 2, 2021, 9:09pm UTC](https://support.prodi.gy/t/prodigy-train-ner-results-explanation/4400 "2021-07-02T21:09:40Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Aziz](https://avatars.discourse-cdn.com/v4/letter/a/ccd318/32.png) [@Aziz](https://support.prodi.gy/u/Aziz)\
**Post date:** [July 2, 2021, 9:09pm UTC](https://support.prodi.gy/t/prodigy-train-ner-results-explanation/4400/1 "2021-07-02T21:09:40Z")

</div>

This is probably a naive question, so after you train the model for ner in prodigy you get the f1 score, recall and precision. My question is is this score calculated on token level or span level for example if the model predict part of the span ( 2 tokens out of 3 tokens that it has to predict), is the whole prediction considered wrong or do you calculate a score ? Thanks in Advance !

---

<div class="post-metadata">

**Author:** ![adriane](https://avatars.discourse-cdn.com/v4/letter/a/46a35a/32.png) [@adriane](https://support.prodi.gy/u/adriane)\
**Post date:** [July 5, 2021, 7:10am UTC](https://support.prodi.gy/t/prodigy-train-ner-results-explanation/4400/2 "2021-07-05T07:10:10Z")

</div>

No, this is a good question and the documentation could be improved here.

It's a micro-PRF on the span level. If any part of the span is wrong (start / end / label), the whole span is counted as wrong.

---

<div class="post-metadata">

**Author:** ![Aziz](https://avatars.discourse-cdn.com/v4/letter/a/ccd318/32.png) [@Aziz](https://support.prodi.gy/u/Aziz)\
**Post date:** [July 5, 2021, 11:03am UTC](https://support.prodi.gy/t/prodigy-train-ner-results-explanation/4400/3 "2021-07-05T11:03:56Z")

</div>

Thanks Adriane

---

<div class="post-metadata">

**Author:** ![Aziz](https://avatars.discourse-cdn.com/v4/letter/a/ccd318/32.png) [@Aziz](https://support.prodi.gy/u/Aziz)\
**Post date:** [July 6, 2021, 6:35pm UTC](https://support.prodi.gy/t/prodigy-train-ner-results-explanation/4400/4 "2021-07-06T18:35:05Z")

</div>

Hey Adriana after doing some manual inspection of the evaluation dataset and the model's predictions, the model's performance on percent shows 100% on precision, recall and f1 score but when I do ner correct on the same evaluation dataset I see that in percent in only grabs the value with the sign (only 10 and not 10 %) I checked the evaluation dataset I m certain that it should label 10% so are you really sure that this performance is calculated on span level ?  
 ![Screen Shot 2021-07-06 at 7.29.56 PM](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/e/e2e42d5616b19efbef6423cb0eaa7e3798e46c19.png)

 ![Screen Shot 2021-07-06 at 3.35.19 PM](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/4/41c1d3e63c99e00443d1a1e4ee6d25b894cbe3e8.jpeg)

---

<div class="post-metadata">

**Author:** ![adriane](https://avatars.discourse-cdn.com/v4/letter/a/46a35a/32.png) [@adriane](https://support.prodi.gy/u/adriane)\
**Post date:** [July 7, 2021, 8:34am UTC](https://support.prodi.gy/t/prodigy-train-ner-results-explanation/4400/5 "2021-07-07T08:34:15Z")

</div>

I am sure that the evaluation is on the span level.

It's hard to know from a distance what's going on with the example above. Can you try running the model in spacy on that exact evaluation text and inspect the annotated entity spans?
