# Automating the annotation for textcat.teach base on score

**URL:** <https://support.prodi.gy/t/automating-the-annotation-for-textcat-teach-base-on-score/64>\
**Category:** Uncategorized\
**Tags:** usage, textcat\
**Created:** [October 25, 2017, 5:51am UTC](https://support.prodi.gy/t/automating-the-annotation-for-textcat-teach-base-on-score/64 "2017-10-25T05:51:33Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![tiru](https://avatars.discourse-cdn.com/v4/letter/t/74df32/32.png) [@tiru](https://support.prodi.gy/u/tiru)\
**Post date:** [October 25, 2017, 5:51am UTC](https://support.prodi.gy/t/automating-the-annotation-for-textcat-teach-base-on-score/64/1 "2017-10-25T05:51:33Z")

</div>

Hi i have 20K samples with 25 different labels, if i want to use prodigy,if i call textcat.teach,it will give web API on which i have to through 20k samples to annotate it,later i can train.

Is there any way to automate this process.? sorry is the query is silly.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [October 25, 2017, 9:30am UTC](https://support.prodi.gy/t/automating-the-annotation-for-textcat-teach-base-on-score/64/2 "2017-10-25T09:30:12Z")

</div>

Just to make sure I understand your question correctly: The 20k examples you have are already labelled, so you want to skip the annotation part?

In this case, you can simply use the [`prodigy db-in` command](https://prodi.gy/docs/recipes#db-in) to import your annotations and add them to a new dataset. All you need to do is convert your data to a format Prodigy can read in – the most convenient would be JSON or JSONL (newline-delimited JSON, which can be read in line by line):

```json
{"text": "Some text", "label": "LABEL"}
{"text": "Some other text", "label": "OTHER_LABEL"}

```

You can then create a new dataset and import the data:

```bash
prodigy dataset my_dataset "Description of my dataset"
prodigy db-in my_dataset /path/to/data.jsonl

```

If no `"answer"` key is present on the examples you’re importing, Prodigy will automatically set them all to `"answer": "accept"` – i.e. import them as correct examples.

---

<div class="post-metadata">

**Author:** ![tiru](https://avatars.discourse-cdn.com/v4/letter/t/74df32/32.png) [@tiru](https://support.prodi.gy/u/tiru)\
**Post date:** [October 25, 2017, 12:00pm UTC](https://support.prodi.gy/t/automating-the-annotation-for-textcat-teach-base-on-score/64/3 "2017-10-25T12:00:42Z")

</div>

Thank you

I have done the same, but accuracy results are very bad on the data which i have imported, for the same data with simple linear svm, i got around 63% accuracy, but with prodigy train not so good results.

**_Note : i am using German news model, i am dealing with German data and it very imbalanced._**

Loaded model de\_dep\_news\_sm  
Using 20% of examples (819) for evaluation  
Using 100% of remaining examples (3278) for training  
Dropout: 0.2 Batch size: 10 Iterations: 5

# LOSS F-SCORE ACCURACY

01 4728.025 0.999 0.999  
02 5579.310 0.999 0.999  
03 6250.201 0.999 0.999  
04 6411.890 0.999 0.999  
05 6350.550 0.999 0.999

MODEL USER COUNT  
accept accept 818  
accept reject 0  
reject reject 0  
reject accept 1

Correct 818  
Incorrect 1

Baseline 1.00  
Precision 1.00  
Recall 1.00  
F-score 1.00  
Accuracy 1.00

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [October 25, 2017, 1:13pm UTC](https://support.prodi.gy/t/automating-the-annotation-for-textcat-teach-base-on-score/64/4 "2017-10-25T13:13:15Z")

</div>

> [@tiru](#):
>
> I have done the same, but accuracy results are very bad on the data which i have imported, for the same data with simple linear svm, i got around 63% accuracy, but with prodigy train not so good results.

Well, according to the results you've posted, Prodigy thinks the accuracy is `1.0`, i.e. 100% – which is obviously suspicious. The reason for this is that you're currently only training on `"accept"` examples, i.e. correct ones. So in this case, your model has simply learned that "everything is correct", which leads to a 100% accuracy, but is obviously pretty useless overall.

The solution is to add a bunch of "wrong" examples – ideally, the same amount as "correct" examples, so you have a nice 50/50 split. You can also do this programmatically by simply swapping out the labels in your existing dataset. (In the long run, you should probably also consider creating wrong examples using different texts.)

Then you add those examples to your Prodigy dataset, and set the answer to `"reject"`:

```bash
prodigy db-in my_dataset /path/to/wrong_data.jsonl --answer reject

```

You should now have 40k annotations in your set – 20k correct and 20k incorrect. Now running `batch-train` should give you a more realistic accuracy score, and hopefully beat your previous 63% 😀

---

<div class="post-metadata">

**Author:** ![tiru](https://avatars.discourse-cdn.com/v4/letter/t/74df32/32.png) [@tiru](https://support.prodi.gy/u/tiru)\
**Post date:** [October 25, 2017, 1:17pm UTC](https://support.prodi.gy/t/automating-the-annotation-for-textcat-teach-base-on-score/64/5 "2017-10-25T13:17:23Z")

</div>

thank you , will try 🙂 , hopefully i will beat the 63% 😉
