# what to do if train-curve shows slight decrease in last sample

**URL:** <https://support.prodi.gy/t/what-to-do-if-train-curve-shows-slight-decrease-in-last-sample/4306>\
**Category:** Uncategorized\
**Tags:** usage, best-practices, training\
**Created:** [June 8, 2021, 7:28am UTC](https://support.prodi.gy/t/what-to-do-if-train-curve-shows-slight-decrease-in-last-sample/4306 "2021-06-08T07:28:46Z")\
**Posts on this page:** 1\
**Showing post:** 7

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [June 8, 2022, 4:57pm UTC](https://support.prodi.gy/t/what-to-do-if-train-curve-shows-slight-decrease-in-last-sample/4306/7 "2022-06-08T16:57:56Z")

</div>

Hi @dave-espinosa!

Excellent questions!

> [@dave-espinosa](#):
>
> **What is the metric obtained in** `train-curve`? **(accuracy, precision, F1...)**

This is from [spaCy scorer](https://spacy.io/api/scorer). You can see the code/formulas from here:

> [@Specific formula for F score, precision and recall NER](https://support.prodi.gy/t/specific-formula-for-f-score-precision-and-recall-ner/4422/2):
>
> Hi! If you're training from manually created annotation, the evaluation all happens within spaCy and doesn't depend on Prodigy. spaCy uses a very standard NER evaluation. If you're working with spaCy v2.x, you can view the code here: For spaCy v3.x, it's here: If you want to do a comparative evaluation, you can also just run both your models over your evaluation data and then calculate the accuracy however you want to, and consistently for both evaluations. Some thing to keep in mind her…

FYI if you're interested, there are ways to add custom evaluation metrics (this uses `textcat` but should be similar for `ner`):

> [@Additional metrics (recall, precision, accuracy F1) in textcat.train-curve](https://support.prodi.gy/t/additional-metrics-recall-precision-accuracy-f1-in-textcat-train-curve/2046):
>
> Hi-- I was wondering whether it is possible at all to include (and do you have plans to include) additional performance metrics in the output of textcat.train-curve. Accuracy is not always the most useful when dealing with unbalanced classes (as I am). Are additional metrics in the pipeline for textcat.train-curve, and do you suggest any workarounds in the meantime? (I guess apart from manually splitting up the data in various sizes and then running textcat.batch-train on them). Cheers!

> [@dave-espinosa](#):
>
> **where are those annotation guidelines located**?

Excellent question! I searched Prodigy/spaCy documentation and found there isn't documentation on creating annotation schemes. I suspect @SofieVL meant annotation schemes in general in terms of carefully defining what each entity means. This made me realize I think there's a lot of potential opportunity with guidelines to help users.

The closest documentation I know of is Matt's 2018 PyData talk. Around 8 minutes into it, he goes through the "[Applied NLP Pyramid of Greatness](https://twitter.com/pmbaumgartner/status/1530160343292465152?s=20&t=O97UZNXMD0566vTl8Tz1Lg)" where he discusses the role of defining an annotation scheme. He then goes through an example of an inadequate annotation scheme for `ner`. Hopefully this will give you a bit of an idea on why annotation schemes can be important.

[![](https://img.youtube.com/vi/jpWqz85F_4Y/maxresdefault.jpg "Building new NLP solutions with spaCy and Prodigy - Matthew Honnibal") ](https://www.youtube.com/watch?v=jpWqz85F_4Y&t=375)

Let me know if you have any thoughts or questions! We greatly appreciate your feedback/questions so please keep them coming 🙂

---

_[View the full topic](https://support.prodi.gy/t/what-to-do-if-train-curve-shows-slight-decrease-in-last-sample/4306)._
