# Ner evaluation probability threshold

**URL:** <https://support.prodi.gy/t/ner-evaluation-probability-threshold/3361>\
**Category:** Uncategorized\
**Tags:** usage, spacy, ner\
**Created:** [September 2, 2020, 7:45am UTC](https://support.prodi.gy/t/ner-evaluation-probability-threshold/3361 "2020-09-02T07:45:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![kak-to-tak](https://avatars.discourse-cdn.com/v4/letter/k/838e76/32.png) [@kak-to-tak](https://support.prodi.gy/u/kak-to-tak)\
**Post date:** [September 2, 2020, 7:45am UTC](https://support.prodi.gy/t/ner-evaluation-probability-threshold/3361/1 "2020-09-02T07:45:39Z")

</div>

I am sorry if it has already been asked, but after running `prodigy train ner ...` the evaluation metrics appear. What I would like to know is under which score do predictions are considered valid during this evaluation.  
Thank you.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [September 4, 2020, 1:44pm UTC](https://support.prodi.gy/t/ner-evaluation-probability-threshold/3361/2 "2020-09-04T13:44:46Z")

</div>

Hi, I'm not 100% sure I understand your question! Prodigy's training command is mostly a thin wrapper around spaCy's API and what you see are the results returned by [`nlp.evaluate`](https://spacy.io/api/language#evaluate). For named entities, that's the precision, recall and f-score.

---

<div class="post-metadata">

**Author:** ![kak-to-tak](https://avatars.discourse-cdn.com/v4/letter/k/838e76/32.png) [@kak-to-tak](https://support.prodi.gy/u/kak-to-tak)\
**Post date:** [September 15, 2020, 7:42am UTC](https://support.prodi.gy/t/ner-evaluation-probability-threshold/3361/3 "2020-09-15T07:42:32Z")

</div>

Thank you for the answer!  
Yes, but on which score do we take an answer as an answer to evaluate?  
I mean, suppose, the model predicts something with a score 0.4, do we take it with a bunch of other answers and evaluate? Or should it be higher? Or can it be lower?
