# How much training data for multiclass/multilabel text classification?

**URL:** <https://support.prodi.gy/t/how-much-training-data-for-multiclass-multilabel-text-classification/6060>\
**Category:** Uncategorized\
**Created:** [October 31, 2022, 5:21am UTC](https://support.prodi.gy/t/how-much-training-data-for-multiclass-multilabel-text-classification/6060 "2022-10-31T05:21:58Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [October 31, 2022, 1:11pm UTC](https://support.prodi.gy/t/how-much-training-data-for-multiclass-multilabel-text-classification/6060/2 "2022-10-31T13:11:11Z")

</div>

hi @joebuckle!

> [@joebuckle](#):
>
> Is there a way for us to know how much more data we need to get a much higher accuracy/score? Or is it more of a trial and error?

Yes! Have you tried the [`train-curve`](https://prodi.gy/docs/recipes#train-curve) recipe?

This recipe is designed to test how accuracy improves with more annotated data. The key is to look for the shape of the training curve. If you find your curve is increasing near the end (e.g., last 25%), this indicates you may get incremental value (information) by labeling more. However, if you find the curve is starting to "level off", this can indicate "diminishing marginal returns". Said differently, there isn't a lot of value of labeling more. If you want to improve your model, you may find you need to rethinking your annotation scheme (e.g., change your class definitions).

You can find several support issues that help in the interpretation and use of `train-curve`:

> [@what to do if train-curve shows slight decrease in last sample](https://support.prodi.gy/t/what-to-do-if-train-curve-shows-slight-decrease-in-last-sample/4306):
>
> I'm creating a new dataset, and so far I made about 500 annotations. I ran the train-curve command, and the score decreased from 0.3 to 0.29 in the last sample. What should I do at this point to make sure I don't annotate a dataset that won't work? Are there some strategies like going back to make sure that the annotation was more consistent, or troubleshoot and find a root cause if possible? Should I just keep annotating and hope that it improves? Thank you!

> [@Best Practices for text classifier annotations](https://support.prodi.gy/t/best-practices-for-text-classifier-annotations/135/2):
>
> Thanks for the questions and sharing your use case! What you’re trying to do definitely sounds feasible, so here are some answers and ideas: In the beginning, you usually want a higher number of accepted examples – there are many thing you might not want your model to learn, so it’s always good to start off with some examples of what you do want. A good way to do this is to start off with a list of seed terms that are very likely to occur in texts your label applies to. You can see an example …

There are also tips on customizing `train-curve`, like adding label stats and saving results:

> [@NER train curve with label stats?](https://support.prodi.gy/t/ner-train-curve-with-label-stats/5052/2):
>
> Hi! At the moment, the per-label stats are only available in the regular training, since it'd otherwise get very verbose very quickly. The per-label stats can also be a bit less representative when training with small portions of the data, because you can easily end up with very few instances of a given label. That said, you can take a look at the implementation in recipes/train.py (you can run prodigy stats to find the location of your Prodigy installation) and make a small adjustment to how t…

Or modifying the evaluation metrics:

> [@\`train-curve textcat\` - display AUC for each classification label](https://support.prodi.gy/t/train-curve-textcat-display-auc-for-each-classification-label/3036):
>
> I would like to see training curves for each classification label using AUC as the metric. It doesn't seem like spacy supports this out of the box, in which case where would I start if I wanted to add reporting functionality that looked something like the below? =============================== sparkles Train curve =============================== % POSITIVE NEGATIVE NEUTRAL ---- -------- -------- ------- 0% 0.40 0.5 0.5 10% 0.88 0.49 0.65 20% 0.86 0.52 0.71 …

> [@joebuckle](#):
>
> We currently have 1000 examples, split into 80-20 training/eval.

Are you creating your own evaluation dataset on your own or are you allowing Prodigy to do it for you automatically?

As the previous post mentions, you may want to create a **dedicated** hold-out (evaluation) dataset if you haven't already. In the `train` docs, there's this tip:

> For each component, you can provide optional datasets for evaluation using the `eval:` prefix, e.g.` --ner dataset,eval:eval_dataset`. If no evaluation sets are specified, the `--eval-split` is used to determine the percentage held back for evaluation.

Let me know if you have any questions on how to write a script for this.

As you may have tried, to improve specific labels, you can use active learning (`textcat-teach`), model-in-the-loop predictions (e.g., `textcat.correct`), or patterns (rules). Here are the [docs](https://prodi.gy/docs/text-classification#active-learning) for doing this with text classification.

Last, if you're working with multiple annotators, another approach to answering " How much data do I need to label?" is to consider bootstrapping for inter-rater reliability. My colleague Peter Baumgartner recently wrote an interesting blog post:

> **["How much data do I need to label?" - The Bootstrapped Inter-Rater...](https://www.peterbaumgartner.com/blog/how-much-data-bootstrap-irr/)**
>
> One of the most frequent questions that arises when doing applied ML projects is “How much data to I need to label?” When I get asked this question, I usually ask a few questions in return: what’s the base rate of the outcome that you’re labeling?...

It's important to note that bootstrapping is a general concept that can be used for any statistic, but is typically computationally intensive which is the limiting factor. For example, you could "bootstrap" (sample with replacement) `train-curve` which would provide uncertainty estimates on accuracy. The problem is this may take a very long time to estimate.

---

_[View the full topic](https://support.prodi.gy/t/how-much-training-data-for-multiclass-multilabel-text-classification/6060)._
