# LLM and bulk annotation

**URL:** <https://support.prodi.gy/t/llm-and-bulk-annotation/6561>\
**Category:** Uncategorized\
**Tags:** ner, spancat\
**Created:** [May 29, 2023, 11:44am UTC](https://support.prodi.gy/t/llm-and-bulk-annotation/6561 "2023-05-29T11:44:52Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![shainaraza](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/shainaraza/32/3813_2.png) [@shainaraza](https://support.prodi.gy/u/shainaraza)\
**Post date:** [May 29, 2023, 11:44am UTC](https://support.prodi.gy/t/llm-and-bulk-annotation/6561/1 "2023-05-29T11:44:52Z")

</div>

if I use LLM to annotate like 1000 samples ( [How can language models augment the annotation process? (ljvmiranda921.github.io)](https://ljvmiranda921.github.io/notebook/2023/03/24/llm-annotation/)) for span categorization, then can I USE THAT 1000 SAMPLES TO ANNOTATE 100000 samples for a span categorization task, what recipe should I use?

---

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [May 31, 2023, 12:03pm UTC](https://support.prodi.gy/t/llm-and-bulk-annotation/6561/2 "2023-05-31T12:03:58Z")

</div>

Hi @shainaraza ,

and welcome to the forum 🙂  
I assume you have a span dataset with 1000 examples and now you're looking for a way to boostrap your annotation further with these examples.  
One way to do that would be to train a small model for predicting spans e.g. spaCy [SpanCategorizer](https://spacy.io/api/spancategorizer) and then use it in Prodigy's [spans.correct](https://prodi.gy/docs/recipes#spans-correct) recipe to streamline the annotation of the big dataset.  
You can find documentation on how to train SpanCategorizer with Prodigy or with spaCy directly [here](https://prodi.gy/docs/span-categorization#training-spacy).

Another way would be to add some representative examples to your prompt and try few-shot annotation (you would use the same recipe that you used for your initial annotation, which I believe is [this](https://github.com/ljvmiranda921/scratch/blob/master/2023-02-16-ukp-argmin/scripts/recipes/spans.py) but you'd use the `--examples_path` option to provide examples. [Here](https://github.com/explosion/prodigy-openai-recipes/tree/main) you can find some documentation on the use of examples and an [example file](https://github.com/explosion/prodigy-openai-recipes/blob/main/examples/ner.yaml). (Lj's recipe you used for spans is very similar to ner recipes documented there and it has the same CLI options.)  
That said, you normally would like to provide just **a few** good examples to make sure the model generalizes from them so you definitely won't be using the entire set 1000.

---

<div class="post-metadata">

**Author:** ![ljvmiranda921](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ljvmiranda921/32/3197_2.png) [@ljvmiranda921](https://support.prodi.gy/u/ljvmiranda921)\
**Post date:** [June 2, 2023, 9:14am UTC](https://support.prodi.gy/t/llm-and-bulk-annotation/6561/3 "2023-06-02T09:14:08Z")

</div>

Hi @shainaraza , just to add to @magdaaniol 's reply, you might also want to curate the 1000 LLM-annotated samples before training a [SpanCategorizer](https://spacy.io/api/spancategorizer) out of it. By that, we mean passing it on to `spans.correct` to build a gold-annotated dataset. While you're in that step, it might also be wise to check if all your span labels are properly represented in those 1000 samples.

Hope that helps!
