# CSV File Text Annotation

**URL:** <https://support.prodi.gy/t/csv-file-text-annotation/2637>\
**Category:** Uncategorized\
**Tags:** usage, solved\
**Created:** [March 8, 2020, 1:06pm UTC](https://support.prodi.gy/t/csv-file-text-annotation/2637 "2020-03-08T13:06:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dinnuv](https://avatars.discourse-cdn.com/v4/letter/d/41988e/32.png) [@dinnuv](https://support.prodi.gy/u/dinnuv)\
**Post date:** [March 8, 2020, 1:06pm UTC](https://support.prodi.gy/t/csv-file-text-annotation/2637/1 "2020-03-08T13:06:28Z")

</div>

Hi,

I have a csv file with feedbacks that I need to annotate using about 8 primary labels and around 25 secondary labels.. I am unable to find the right documentation for it. Can you help me with this?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 9, 2020, 10:23am UTC](https://support.prodi.gy/t/csv-file-text-annotation/2637/2 "2020-03-09T10:23:07Z")

</div>

Hi! You can either read in your CSV file directly or convert it to JSONL. Check out this docs section for the expected file formats: [https://prodi.gy/docs/api-loaders#input](https://prodi.gy/docs/api-loaders#input)

How you do the annotation depends on what exactly you need to label. If you need to assign labels to the whole texts, check out the docs on [text classification](https://prodi.gy/docs/text-classification). Also see this section for tips on how to efficiently handle large and/or hierarchical label schemes, like your primary and secondary labels: [https://prodi.gy/docs/text-classification#large-label-sets](https://prodi.gy/docs/text-classification#large-label-sets)

---

<div class="post-metadata">

**Author:** ![dinnuv](https://avatars.discourse-cdn.com/v4/letter/d/41988e/32.png) [@dinnuv](https://support.prodi.gy/u/dinnuv)\
**Post date:** [March 9, 2020, 7:02pm UTC](https://support.prodi.gy/t/csv-file-text-annotation/2637/3 "2020-03-09T19:02:05Z")

</div>

Hi Ines, can you give me an example. I am really struggling to make this work... I am trying to annotate feedbacks in a csv file which has 2 columns, an "ID" and "Feedback". I tried the below recipe so that I can get "ID" in my JSONL output. But, when I run the command, I am seeing error saying:

**[x] Error while validating stream: no first example**  
**This likely means that your stream is empty.**

I ran the following command:  
**python -m prodigy feedback\_recipe dataset "C:/Users/...../feedback\_prodigy\_test.csv" -F temp.py**

**Below is the recipe:**

1.) From where should I run the command?  
2.) Can you explain the details regarding "dataset" in the command  
3.) DIs the command that I ran correct?

import csv  
import prodigy

@prodigy.recipe('feedback\_recipe',  
dataset=prodigy.recipe\_args['dataset'],  
file\_path=("C:/Users/......../feedback\_prodigy\_test.csv", "positional", None, str))

def feedback\_recipe(dataset, file\_path):  
"""Annotate the feedbacks using different labels."""  
stream = custom\_csv\_loader(file\_path) # load in the CSV file  
stream = add\_options(stream) # add options to each task

```
return {
      'dataset': dataset, # save annotations in this dataset
      'view_id': 'choice', # use the choice interface
      'config': {'choice_style': 'multiple'},
      'stream':stream,
  }

```

def custom\_csv\_loader(file\_path):  
with open(file\_path) as csvfile:  
reader = csv.DictReader(csvfile)  
for row in reader:  
id = row.get('ResponseId')  
text = row.get('Feedback\_Explanation')  
yield {'text': text, 'meta': {'id':id}}

def add\_options(stream):  
# Helper function to add options to every task in a stream  
options = [  
{"id": "Billing & Payment", "text": "Billing & Payment"},  
{"id": "Registration & Sign-In", "text": "Registration & Sign-In"},  
{"id": "Website Issues", "text": "Website Issues"},  
{"id": "Customer Service", "text": "Customer Service"},  
]  
for task in stream:  
task["options"] = options  
yield task

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 11, 2020, 9:02am UTC](https://support.prodi.gy/t/csv-file-text-annotation/2637/4 "2020-03-11T09:02:00Z")

</div>

> [@dinnuv](#):
>
> I am trying to annotate feedbacks in a csv file which has 2 columns, an "ID" and "Feedback".

If you're trying to load a CSV file "as is", the text should be in a column `text` or `Text`. Otherwise, Prodigy can't know where to look for it. Check out the link I posted above for examples of the data format: [Loaders and Input Data · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/api-loaders#input)

You shouldn't really need a custom recipe for what you're trying to do – `textcat.manual` should do all you need? Check out the documentation on text classification here:

> **[Text Classification · Prodigy · An annotation tool for AI, Machine Learning...](https://prodi.gy/docs/text-classification/)**
>
> A downloadable annotation tool for NLP and computer vision tasks such as named entity recognition, text classification, object detection, image segmentation, A/B evaluation and more.

> [@dinnuv](#):
>
> 1.) From where should I run the command?  
> 2.) Can you explain the details regarding "dataset" in the command  
> 3.) DIs the command that I ran correct?

You might want to check out the "Prodigy 101" guide, which explains how to get started, how to run commands and what the arguments mean and how to set up your annotation projects. It also has a glossary at the end that explains the most common terms, like "dataset" etc.

> **[Prodigy 101 – everything you need to know · Prodigy · An annotation tool for...](https://prodi.gy/docs/)**
>
> A downloadable annotation tool for NLP and computer vision tasks such as named entity recognition, text classification, object detection, image segmentation, A/B evaluation and more.
