# Annotated Dataset and NER task with Prodigy

**URL:** <https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155>\
**Category:** Uncategorized\
**Tags:** usage, ner\
**Created:** [January 4, 2022, 9:14pm UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155 "2022-01-04T21:14:04Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![sudarshan85](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sudarshan85/32/2199_2.png) [@sudarshan85](https://support.prodi.gy/u/sudarshan85)\
**Post date:** [January 4, 2022, 9:14pm UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155/1 "2022-01-04T21:14:04Z")

</div>

Hello,

I'm new to using Prodigy and have a few questions. Our team is working on a task to extract detoxification events from clinical notes written by healthcare providers. We define our custom label for each detox event and we approaching this as a named entity recognition task for automatically labeling certain events based on SME guidance.

Currently, we are working with 15095 snippets. We've had a SME use Prodigy with the following command:

```python
prodigy ner.manual detox_event_extraction blank:en data_detox.jsonl --label <labels>

```

The actual labels we use is not important. The SME has annotated 1000 snippets out of the 15095 snippets and saved it a database file called `detox_event_extraction.db`. I've extracted the those 1000 annotated snippets in a separate json file `annotated_snippets.jsonl`. I have the following questions:

1. What information does the database actually hold? Specifically, does it hold only annotated dataset or does it hold the entire dataset?
2. Due to some issues, I had to delete the database file. But I have the 1000 annotated snippets (`annotated_snippets.jsonl`) and the original data file (`data_detox.jsonl`). How do I recreate the database file using these files so that if needed the SME can continue annotating from snippet 1001?

The workflow that I'd like to follow is similar to the ingredients NER [video](https://www.youtube.com/watch?v=59BKHO_xBPA&t=1263s) by Ines Montani. Specifically, I want to use the 750 out of the 1000 annotated snippets to train a spacy model (250 for eval) and then use a `ner.correct` recipe with the SME to annotate a further 1000 notes. Eventually, I want to use `ner.teach` for active learning with the SME to train the model with another few 1000 notes. I'm saving a final 5000+ snippets as a final test set. I have the following questions:

1. Is what I have described a typical workflow for the task that I want done?
2. I'm having a hard time understanding how the database file incorporates all the annotations provided by `ner.manual`, `ner.correct`, and `ner.teach`. Do I save different database files for each session?
3. Does the model continuously update with `ner.correct`? If not, do I just train a new model with updated annotations?
4. Does the model continuously update with `ner.teach`? Is there an example how I could go about doing this?

I apologize for the long post and would appreciate any help regarding thes questions.

Thanks!

---

<div class="post-metadata">

**Author:** ![ljvmiranda921](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ljvmiranda921/32/3197_2.png) [@ljvmiranda921](https://support.prodi.gy/u/ljvmiranda921)\
**Post date:** [January 6, 2022, 1:43am UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155/2 "2022-01-06T01:43:00Z")

</div>

Hi @sudarshan !

> [@sudarshan85](#):
>
> What information does the database actually hold? Specifically, does it hold only annotated dataset or does it hold the entire dataset?

It only holds the annotated dataset. Whenever you're done annotating and once you "saved" your annotations, their values (whether you accepted, rejected, or ignored them) will be saved into the db.

> [@sudarshan85](#):
>
> Due to some issues, I had to delete the database file. But I have the 1000 annotated snippets (`annotated_snippets.jsonl`) and the original data file (`data_detox.jsonl`). How do I recreate the database file using these files so that if needed the SME can continue annotating from snippet 1001?

You can use the `db-in` command. Prodigy will skip annotations that were already in the dataset and you can just keep annotating with the original file 🙂

---

<div class="post-metadata">

**Author:** ![ljvmiranda921](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ljvmiranda921/32/3197_2.png) [@ljvmiranda921](https://support.prodi.gy/u/ljvmiranda921)\
**Post date:** [January 6, 2022, 1:46am UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155/3 "2022-01-06T01:46:50Z")

</div>

Oops my reply got cut, here are the answers for the next set of questions 🙂 Thanks for listing them in an accessible manner, anyway! 😃

> [@sudarshan85](#):
>
> Is what I have described a typical workflow for the task that I want done?

The typical workflow is to annotate everything, then do a final training with all of your annotated data. Ideally, we'd only use `ner.teach` to improve the quality-of-life of our annotation process, we won't use it to train a model that goes to prod.

Your workflow is OK if you're just going to annotate, but in the end, you'd want to use all those annotations from ner.correct and ner.teach to train a final model (which you can conveniently do via `prodigy train` ).

> [@sudarshan85](#):
>
> I'm having a hard time understanding how the database file incorporates all the annotations provided by `ner.manual`, `ner.correct`, and `ner.teach`. Do I save different database files for each session?

There's only one database and one database file, but within that you have datasets (collection of annotations). Typically, you'd use different datasets for different annotation experiments and types (`manual`, `binary`, etc.). You can then train from multiple datasets later on.

> [@sudarshan85](#):
>
> Does the model continuously update with `ner.correct`? If not, do I just train a new model with updated annotations?

You can check the [`--update` parameter for ner.correct](https://prodi.gy/docs/recipes#ner-correct), it gives you the option to update your model in the loop. But in the end, yes, you'd still want to train a new model (from scratch) using the updated annotations.

> [@sudarshan85](#):
>
> Does the model continuously update with `ner.teach`? Is there an example how I could go about doing this?

Yes it does update, but it's better if you train again from scratch (with the updated annotations) for your production/final model. WIth that, you get all the conveniences of setting training hyperparameters, refining your config, etc. etc.

---

<div class="post-metadata">

**Author:** ![sudarshan85](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sudarshan85/32/2199_2.png) [@sudarshan85](https://support.prodi.gy/u/sudarshan85)\
**Post date:** [January 6, 2022, 4:23pm UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155/4 "2022-01-06T16:23:24Z")

</div>

Hi Lj Miranda,

Thank you for your very thorough answer. I will try out your suggestions and report back with any questions I come up with.

While I understand the workflow that you describe, I have a non-typical situation where the model wouldn't really go into production. This is a proof-of-concept presentation of Prodigy and its capabilities in a very specific clinical concept setting. In the end, my "product" will be detailed instructions tailored to the specifications of a task so that SME's with minimal expertise in ML and NLP maybe be able to use Prodigy for their task.

I'm still have having a hard time understanding the difference between _database_ and _dataset_. In the video I linked, Ines recommends saving different databases for different parts of the workflow. In that case, will each database have different datasets? Forgive me for my ignorance on this issue.

---

<div class="post-metadata">

**Author:** ![ljvmiranda921](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ljvmiranda921/32/3197_2.png) [@ljvmiranda921](https://support.prodi.gy/u/ljvmiranda921)\
**Post date:** [January 7, 2022, 1:01am UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155/5 "2022-01-07T01:01:39Z")

</div>

Hi @sudarshan

> [@sudarshan85](#):
>
> I'm still have having a hard time understanding the difference between _database_ and _dataset_.

- You can think of a **dataset** as a "collection of annotations." If you're using Prodigy with its default parameters / configuration, all your datasets will just reside in a single database.
- A **database** , on the other hand, can be liken to a central storage of your datasets. This can be a SQLite, MySQL, etc. database based on your config. They hold all your datasets by default. It's a one-to-many relationship 🙂

---

<div class="post-metadata">

**Author:** ![sudarshan85](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sudarshan85/32/2199_2.png) [@sudarshan85](https://support.prodi.gy/u/sudarshan85)\
**Post date:** [February 4, 2022, 5:43pm UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155/6 "2022-02-04T17:43:11Z")

</div>

I'm posting an update here for this task along with a few questions based on some hickups I've been having.  
As mentioned in the first post this is a NER task to extract clinical concepts (specifically concepts relating to detox) from medical notes. I can't share the note details here but I'll share relevant task details. I used the video example I [linked](https://youtu.be/59BKHO_xBPA) earlier to guide my process.

`detox_data.jsonl` -- The original file containing the clinical notes, this contains 15,895 lines corresponding to that many snippets  
`detox_event_extraction` -- Dataset in the database that was created when SME manually annotated the snippets

1. I launched Prodigy with `ner.manual` with `detox_data.jsonl` and a `blank:en` model to have the SME annotate 1000 documents (over multiple sessions) and save it to the dataset `detox_event_extraction` using the following command:

```python
prodigy ner.manual detox_event_extraction blank:en detox_data.jsonl --label <labels>

```

1. I extracted the annotated dataset to a new `jsonl` file using the command:

```python
prodigy db-out detox_event_extraction > annotated_snippets.jsonl

```

1. For reasons that are not important to this topic, I had to drop the `detox_event_extraction` dataset and I also ended up deleting the database file `detox_event_extraction.db` file. After some messing around and with the help of of the post [here](https://support.prodi.gy/t/editing-datasets/66/2), I was able to get another database file a re-added the annotated snippets to a new dataset called `dee_manual_1000`
2. I then trained a NER model using the _scispacy_ model as a starting point:

```python
prodigy train ner dee_manual_1000 en_core_sci_lg --output dee_manual_1000_model --eval-split 0.2

```

The training took a while and the results were not very good with a F1 score of only 20.  
5. I ran the `train-curve` command which showed good improvements as more data was added. So my next step is to have the SME annotate more documents and iteratively train and annotate to get an acceptable performance.

According the tutorial video, my next step was to launch Prodigy with `ner.correct`:

```python
prodigy ner.correct dee_correct_2000 dee_manual_1000_model detox_data.jsonl --label <labels> --exclude dee_manual_1000

```

I would like to point out three things:

1. I want save the new annotations in a separate dataset `dee_correct_2000` as suggested by the tutorial
2. I'm using the already trained model `dee_manual_1000_model`
3. I would like to exclude the already annotated snippets, hence I've added the `--exclude` option with the appropriate dataset name

This is the point I'm running into couple of issues and I have the following questions:

1. `ner.correct` seems to be doing sentence segmentation and showing only one sentence at a time. I would like it to show the whole document as it would in `ner.manual`. How do I achieve this?
2. Despite adding `--exclude` option in the original command, I see samples from the original annotations which gives a feeling of "starting from scratch" and unfortunately SME time is valuable. Why is this happening?
3. What is the sequence of documents that is displayed when running `ner.manual` vs `ner.correct`? Do they follow the same sequence as presented in the `jsonl` file. I know that `ner.manual` does, but not sure about `ner.correct`.

Thank for your reading this long post and any help that is provided!

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [February 3, 2023, 4:37pm UTC](https://support.prodi.gy/t/annotated-dataset-and-ner-task-with-prodigy/5155/7 "2023-02-03T16:37:36Z")

</div>

hi @sudarshan85!

Sorry for the delay. We're trying to close out old tickets.

> [@sudarshan85](#):
>
> `ner.correct` seems to be doing sentence segmentation and showing only one sentence at a time. I would like it to show the whole document as it would in `ner.manual`. How do I achieve this?

By default, [`ner.correct`](https://prodi.gy/docs/recipes#ner-correct) does sentence segmentation (unlike `ner.manual`. You can turn it off by adding `--unsegmented`.

> [@sudarshan85](#):
>
> Despite adding `--exclude` option in the original command, I see samples from the original annotations which gives a feeling of "starting from scratch" and unfortunately SME time is valuable. Why is this happening?

That's tough to confirm. Let me go through a reproducible example of what should happen.

Start with this source file:  
[nyt\_text\_dedup.jsonl](https://support.prodi.gy/uploads/short-url/bWDH7T6uTm63MU6pA1crkOhu5Zn.jsonl) (18.5 KB)

### Step 1: Label 10 records into dataset `ner_correct1`

```terminal
python -m prodigy ner.correct ner_correct1 en_core_web_sm nyt_text_dedup.jsonl --label LOC

```

I then labeled the first 10 records. You can see them by running:

```terminal
$ python -m prodigy print-dataset ner_correct1

```

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/9/9657aea875b8dbf2d058067c8d4f611ffbee0af8.png)

### Step 2: Rerun but use `--exclude` to exclude records in `ner_correct1`

```terminal
python3 -m prodigy ner.correct ner_correct2 en_core_web_sm data/nyt_text_dedup.jsonl --exclude ner_correct1 --label LOC

```

 ![localhost_8080_ (20)](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/2/23d976c2defb08970801ff2be314f96997a5f201.png)

Notice it starts on record 10 (see metadata in bottom right). Therefore, it skipped the first 10 records.

> [@sudarshan85](#):
>
> What is the sequence of documents that is displayed when running `ner.manual` vs `ner.correct`? Do they follow the same sequence as presented in the `jsonl` file. I know that `ner.manual` does, but not sure about `ner.correct`.

Yes. `ner.manual` and `ner.correct` will use based on order of documents. This is different than [`ner.teach`](https://prodi.gy/docs/recipes#ner-teach), which uses active learning and will alter the order of the documents based on uncertainty scoring.

Let us know if you have any other questions!
