# Does Prodigy supports hierarchical annotation?

**URL:** <https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249>\
**Category:** Uncategorized\
**Tags:** usage\
**Created:** [February 27, 2019, 7:56pm UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249 "2019-02-27T19:56:40Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![nawshad](https://avatars.discourse-cdn.com/v4/letter/n/4491bb/32.png) [@nawshad](https://support.prodi.gy/u/nawshad)\
**Post date:** [February 27, 2019, 7:56pm UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/1 "2019-02-27T19:56:40Z")

</div>

I would like to use Prodigy for hierarchical annotation. For example, if an example qualifies for annotation A, then annotate it and move on to next example, otherwise if it qualifies for B, then check if its also C, D, or E.

Or is there a way to customize prodigy for this kind of task.

Thanks,

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 28, 2019, 8:27am UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/2 "2019-02-28T08:27:58Z")

</div>

Prodigy itself is pretty agnostic to what your annotations “mean” so you can definitely build a workflow like this. The only thing that’s kinda built-in is a strong focus on single decisions at a time and automating as much as possible.

One approach could be to use the [choice interface](https://prodi.gy/demo?view_id=textchoice) and start by annotating the top-level buckets like A, B and C, without worrying about the lower-level categories. [See here](https://prodi.gy/docs/workflow-custom-recipes#example-choice) for an example recipe code. In the next step, you can then stream in the examples again and add different options, based on the top-level category that was selected – for example, A1 and A2 for A and so on. Prodigy streams are regular Python generators, so you can automate all of this logic by putting it in a function that yields annotation examples. For example, something like this:

```python
hierarchy = {'A': ['A1', 'A2'], 'B': ['B1', 'B2'], 'C': ['C1', 'C2']}

def get_stream(examples):
    for eg in examples: # the examples with top-level categories
        top_labels = eg['accepted'] # ['A'] or ['B', 'C'] if multiple choice
        for label in top_labels:
            sub_labels = hierarchy[label]
            options = [{'id': opt, 'text': opt} for opt in sub_labels]
            # create new example with text and sub labels as options
            new_eg = {'text': eg['text'], 'options': options}
            yield eg

```

Doing the levels in separate steps also allows you to iterate faster if you end up having to adjust the annotation scheme. Not all schemes are set in stone and if your annotators struggle with a top-level decision like B vs. C, they’ll likely also struggle with the lower-level decisions. So ideally, you want to find out about this as early as possible and before you commission the full fine-grained annotations on your entire corpus.

---

<div class="post-metadata">

**Author:** ![ks1996](https://avatars.discourse-cdn.com/v4/letter/k/c89c15/32.png) [@ks1996](https://support.prodi.gy/u/ks1996)\
**Post date:** [April 3, 2020, 6:13pm UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/4 "2020-04-03T18:13:31Z")

</div>

Hey Ines, @ines @MatthewC

I am trying to set up a hierarchical classification for classes and sub-classes. I did the first part of the method mentioned above to do the high level classification. I do not understand how the next step to provide next level hierarchy fits in with the first step.

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/6/6024e0c6125d0af6fe636ff3c971519c1c575e14.png)

Is the original data set required to be passed here? I am not sure what you mean by examples with top-level categories.

It would be really helpful if you could explain how the recipe for hierarchy works together with both the steps together?

I am getting the following error when i try to write one single recipe for entire process

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/2/2b20940d25bc73e8ae28b9bcbf56e6641c8908aa.png)

I am trying to class first as fluid and mechanical and later as f1,f2 and m1,m2.  
This is the snippet i used. ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/d/d73a674f838d19f5261fd5043078f94be182791a.png)

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [April 6, 2020, 9:17am UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/5 "2020-04-06T09:17:26Z")

</div>

> [@ks1996](#):
>
> Is the original data set required to be passed here? I am not sure what you mean by examples with top-level categories.

Yes, those are supposed to be the examples you've previously annotated with the top-level categories (e.g. using a recipe like [`textcat.manual`](https://prodi.gy/docs/recipes#textcat-manual) or any other [custom recipe](https://prodi.gy/docs/custom-recipes) with the `choice` interface).

> [@ks1996](#):
>
> I am getting the following error when i try to write one single recipe for entire process

The `eg["accept"]` is referring to the `"accept"` key of the dictionary here, so you shouldn't modify that one. Its value is a list of labels that were selected in the UI. This is the format produced by the choice recipe – see here for an example: [Annotation interfaces · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/api-interfaces#choice)

So the top labels are coming from the data you previously annotated with those labels. And then for each of those examples, you create a new task with the lower-level labels as options. Also see here for a visual example of the concept: [Text Classification · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/text-classification#large-label-sets)

---

<div class="post-metadata">

**Author:** ![ks1996](https://avatars.discourse-cdn.com/v4/letter/k/c89c15/32.png) [@ks1996](https://support.prodi.gy/u/ks1996)\
**Post date:** [April 6, 2020, 3:53pm UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/6 "2020-04-06T15:53:30Z")

</div>

@ines  
This is the custom recipe I am using which will clean and annotate the data at the same time.

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/e/e4fc20f029a71127824384ea6ea9d6041f4e5f85.png)

after selecting the options in the UI and annotating i try converting it to a json file using to-patterns command : python -m prodigy terms.to-patterns examples\_eg hp.jsonl --label fluid,mechanical --spacy-model blank:en

This is the file i get. It is supposed to be either fluid or mechanical and not both ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/8/8558af6440ab56d62c7d518eb0f98c2cad8c42b7.png)

I think this is the reason I am not able to set the hierarchy in the next step. Any thoughts on this?

My recipe for hierarchy is as follows ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/a/a8b0ff2d0455afde1f1360db65f2a4772e4a1093.png)

I am getting the same error as I mentioned in my previous comment.  
 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/d/d259faa8452bf301264ec6da90efe74e683f4b73.png)

I might be missing something here.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [April 7, 2020, 9:17am UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/7 "2020-04-07T09:17:47Z")

</div>

> [@ks1996](#):
>
> after selecting the options in the UI and annotating i try converting it to a json file using to-patterns command : python -m prodigy terms.to-patterns examples\_eg hp.jsonl --label fluid,mechanical --spacy-model blank:en

Why are you converting to [match patterns](https://prodi.gy/docs/named-entity-recognition#manual-patterns) here? I don't think that's what you want to do – I think you just want to export the annotated examples? You can use the `db-out` command, or even load the annotations from the database [programmatically in your recipe](https://support.prodi.gy/t/programmatic-way-to-get-stats/2730/2).

> [@ks1996](#):
>
> I am getting the same error as I mentioned in my previous comment.

This means that `eg` is a string. I think you're missing the step that actually loads the examples? So whatever you pass in as `examples` (like the path to a file) is passed through here. So it's trying to acccess the index `['accept']` of a string like `/path/to/something`, which isn't going to work.

---

<div class="post-metadata">

**Author:** ![ks1996](https://avatars.discourse-cdn.com/v4/letter/k/c89c15/32.png) [@ks1996](https://support.prodi.gy/u/ks1996)\
**Post date:** [April 7, 2020, 5:11pm UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/8 "2020-04-07T17:11:22Z")

</div>

@ines Thank you so much for responding. I really appreciate it.

> [@ines](#):
>
> Why are you converting to [match patterns](https://prodi.gy/docs/named-entity-recognition#manual-patterns) here? I don't think that's what you want to do – I think you just want to export the annotated examples? You can use the `db-out` command, or even load the annotations from the database [programmatically in your recipe](https://support.prodi.gy/t/programmatic-way-to-get-stats/2730/2).

It makes sense now. Thank you.

This is where I am having trouble catching up.

> [@ines](#):
>
> I think you're missing the step that actually loads the examples?

I think I have got everything right till the part where I classify on a high level using custom recipe with choice interface into "fluid" and "mechanical". But later classifying fluid as f1 & f2 and mechanical as m1 & m2 is what I am having trouble.

I am not quite sure of how to load the data to set up the hierarchy. Do you think I should set a generator function which takes the examples line by line and indexes the ['accept'].  
I apologize for taking up your time but I am trying to understand the underlying workflow here to set up a hierarchy custom recipe. It would be really helpful if you could direct me towards an example of such workflow setup.

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [April 7, 2020, 10:04pm UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/9 "2020-04-07T22:04:25Z")

</div>

So the workflow I was proposing would look something like this:

1. Annotate some examples with `textcat.manual` the top-level categories, e.g. "fluid" and "mechanical".
2. Load the data created in the first step in your recipe and create new questions for each example. For instance, for every example you've annotated with "fluid", create a new question that now has the options "f1" and "f2".
3. Annotate again and you have a dataset with all top-level categories and lower-level categories.

So if you've saved your annotations from step 1 in a dataset called `textcat_top_level`, you can run `prodigy db-out textcat_top_level ./output` to save a file `textcat_top_level.jsonl` to the directory `output`. You can then use that as the input in your recipe and use the [`JSONL` loader](https://prodi.gy/docs/api-loaders#loaders-file) to load the examples.

(For some background on custom recipes, you might also find [my video here](https://prodi.gy/docs/custom-recipes#video-imagecaptioning) useful. It's a pretty different topic but I'm also trying to explain the overall concept of recipe scripts and how the pieces fit together.)

---

<div class="post-metadata">

**Author:** ![ks1996](https://avatars.discourse-cdn.com/v4/letter/k/c89c15/32.png) [@ks1996](https://support.prodi.gy/u/ks1996)\
**Post date:** [April 8, 2020, 2:04pm UTC](https://support.prodi.gy/t/does-prodigy-supports-hierarchical-annotation/1249/10 "2020-04-08T14:04:29Z")

</div>

@ines  
Thank you very much for your responses. It does clear a lot of questions for me.
