# ner.correct: Only 31 annotations to database no matter how many actually annotated everytime

**URL:** <https://support.prodi.gy/t/ner-correct-only-31-annotations-to-database-no-matter-how-many-actually-annotated-everytime/3986>\
**Category:** Uncategorized\
**Tags:** ner, database\
**Created:** [March 7, 2021, 11:23am UTC](https://support.prodi.gy/t/ner-correct-only-31-annotations-to-database-no-matter-how-many-actually-annotated-everytime/3986 "2021-03-07T11:23:33Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![AntiLibrary5](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/antilibrary5/32/2004_2.png) [@AntiLibrary5](https://support.prodi.gy/u/AntiLibrary5)\
**Post date:** [March 7, 2021, 11:23am UTC](https://support.prodi.gy/t/ner-correct-only-31-annotations-to-database-no-matter-how-many-actually-annotated-everytime/3986/1 "2021-03-07T11:23:33Z")

</div>

Python: 3.7  
Prodigy: 1.10.6

Hi,  
I followed the guide described in [sense2vec reloaded: contextually-keyed word vectors · Explosion](https://explosion.ai/blog/sense2vec-reloaded) and [https://youtu.be/59BKHO\_xBPA](https://youtu.be/59BKHO_xBPA) for **NER** in patent data from [link](https://www.uspto.gov/web/patents/classification/cpc/html/cpc-H01M.html).

After training a first model,  
python -m prodigy train ner annotatedh01m en\_vectors\_web\_lg --init-tok2vec ./tok2vec\_cd8\_model289.bin --output ./tmp\_model --eval-split 0.2

Moving onto the step where we label more examples by correcting the model's predictions, I worked through 200 examples ( please see the screenshot below) using ner.correct but the output was that only 31 were saved:

python -m prodigy ner.correct annotatedg06f\_correct ./tmp\_model g06fsents.3000.txt --loader txt --label TECH --exclude annotatedg06f

Using 1 label(s): TECH  
Added dataset annotatedg06f\_correct to database SQLite.  
⚠ The model you're using isn't setting sentence boundaries (e.g. via the parser or sentencizer). This means that incoming examples won't be split into sentences.  
✨ Starting the web server at [http://localhost:8080](http://localhost:8080/) ...  
Open the app in your browser and start annotating! ^C

✔ Saved 31 annotations to database SQLite  
Dataset: annotatedg06f\_correct  
Session ID: 2021-02-17\_16-56-43

Using a the older ner.make\_gold had the same output:

python -m prodigy ner.make-gold annotatedh01m\_correct ./tmp\_model h01msents.4802.txt --loader txt --label BATT --exclude annotatedh01m

✔ Saved 31 annotations to database SQLite  
Dataset: annotatedh01m\_correct  
Session ID: 2021-03-07\_11-42-12

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/a/a7ecd5db32007f42408ac3168ff48561176d4356.png)

I have once tried setting up my venv again with the same results and do not understand what the problem might be. Any help is appreciated.  
Thank You.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 8, 2021, 9:56am UTC](https://support.prodi.gy/t/ner-correct-only-31-annotations-to-database-no-matter-how-many-actually-annotated-everytime/3986/2 "2021-03-08T09:56:34Z")

</div>

Hi! That's strange, I haven't seen this one before. If you run `prodigy stats` or access the database in Python, how many examples does it show you are in the datasets? Are you using any custom SQLite database? And did you always hit "save" at the end of the annotation session?

---

<div class="post-metadata">

**Author:** ![AntiLibrary5](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/antilibrary5/32/2004_2.png) [@AntiLibrary5](https://support.prodi.gy/u/AntiLibrary5)\
**Post date:** [March 8, 2021, 2:13pm UTC](https://support.prodi.gy/t/ner-correct-only-31-annotations-to-database-no-matter-how-many-actually-annotated-everytime/3986/3 "2021-03-08T14:13:13Z")

</div>

Hi,  
Thank you for your reponse. Yes I did save everytime. I was guessing that maybe the same examples do not get written resulting in a lower number which makes sense but in multiple sessions I did go through many different examples and corrected them, but saw the output that 31 annotations were saved so it cannot be a coincidence. I re-did 100 annotations just now with ner.correct recipe:

Checking the db directly, at first I have 83 annotations up to now:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/e/e5afbb8aee34f564456dd5484d4f4e2c3d808141.png)

Next I annotate some more (100 in one session):

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/3/319c294ba4bb8db4ca10a701787a84d13559f939.png)

Then after ending the session I see 31 annotations saved:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/0/0044f989b7711676d7de82c4ae769614162c14f7.png)

And prodigy db-out counts a total 114 annotations now:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/6/6f19cdaeaada7bfbf5728768354dfa12821c1a6f.png)

So it seems each time only 31 or less annotations are added via ner.correct recipe.

Please let me know if I'm making some mistake.

Thank you.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 9, 2021, 1:06am UTC](https://support.prodi.gy/t/ner-correct-only-31-annotations-to-database-no-matter-how-many-actually-annotated-everytime/3986/4 "2021-03-09T01:06:17Z")

</div>

Yeah, the 31 definitely makes it very strange 🤔

Under the hood, Prodigy uses `peewee` to manage the database connection. The "saved annotations" message is only shown if the database reports that the data was successfully added. So it's unlikely that the databaese connection is broken. Prodigy will also save batches of data in the background as you annotate and if that fails, you'll see an error. This mechanism is the same for all workflows, so I don't think the problem is recipe-specific.

One thing you could try to help get to the bottom of this: if you run Prodigy with the environment variable `PRODIGY_LOGGING=basic`, you'll see log statements of everything that's going on, including examples saved to the database. If you click through the examples, you should see log statements for the `/give_answers` endpoint and the controller receiving answers as Prodigy auto-saves in the background. What do those logs say?

Also, can you reproduce this with a new dataset? (You can just click through some examples quickly.) If it turns out that only 31 annotations are saved, are these examples from the start or the end of the dataset? Do you see any pattern here?
