# Annotation JSON

**URL:** <https://support.prodi.gy/t/annotation-json/5746>\
**Category:** Uncategorized\
**Tags:** ner, spancat\
**Created:** [June 27, 2022, 4:01pm UTC](https://support.prodi.gy/t/annotation-json/5746 "2022-06-27T16:01:43Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![NNN](https://avatars.discourse-cdn.com/v4/letter/n/58f4c7/32.png) [@NNN](https://support.prodi.gy/u/NNN)\
**Post date:** [June 27, 2022, 4:01pm UTC](https://support.prodi.gy/t/annotation-json/5746/1 "2022-06-27T16:01:43Z")

</div>

Hi,

I'm getting a weird JSON back. For some of the `spans` there are additional keys (`text`, `source` and `_input_hash`) whereas for others these do not appear.

Regarding the `_input_hash`, I'm not sure what value I could gain from it and why is it exactly the same `_input_hash` as the one under the `meta` key.

Regarding the `_task_hash`, I see it can be repeated. I'm also not so sure how to gain value from it.

Regarding the `timestamp`, how exactly is it interpreted? What time units does it show?

I'm attaching a snippet of an annotation JSON

Thank you!

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/f/f684dd30d53671e464770a076b69a008d118941e.png)

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [June 27, 2022, 4:32pm UTC](https://support.prodi.gy/t/annotation-json/5746/2 "2022-06-27T16:32:36Z")

</div>

hi @NNN!

Have you seen the [Prodigy documentation on `_input_hash` and `_task_hash`](https://prodi.gy/docs/api-loaders/#hashing)?

These hashes are used to identify deduplication (e.g., whether two examples are entirely different, different questions about the same input, e.g. text, or the same question about the same input).

Here are details about each hash:

| Hash | Type | Description |
| --- | --- | --- |
| \_input\_hash | int | Hash representing the input that annotations are collected on, e.g. the text, image or html. Examples with the same text will receive the same input hash. |
| \_task\_hash | int | Hash representing the “question” about the input, i.e. the label, spans or options. Examples with the same text but different label suggestions or options will receive the same input hash, but different task hashes. |

Prodigy uses these behind the scene to account for deduplications, so you can ignore them (however they can be helpful in tracking down the road).

> [@NNN](#):
>
> Regarding the `timestamp`, how exactly is it interpreted? What time units does it show?

This is a [unixtime stamp](https://www.unixtimestamp.com/). There are [python converters](https://stackoverflow.com/questions/3682748/converting-unix-timestamp-string-to-readable-date) that can help in converting this to time readable formats.

Let me know if this answers your questions or if you have any further questions!

---

<div class="post-metadata">

**Author:** ![NNN](https://avatars.discourse-cdn.com/v4/letter/n/58f4c7/32.png) [@NNN](https://support.prodi.gy/u/NNN)\
**Post date:** [August 7, 2022, 12:40pm UTC](https://support.prodi.gy/t/annotation-json/5746/3 "2022-08-07T12:40:42Z")

</div>

Hi Ryan,

Thanks so much for your answer. I apologise for my late response.

1.) Hashes - **understood**.

2.) **Unclear** - I still don't understand why for some of the `spans` there are additional keys (`text` , `source` and `_input_hash` ) whereas for others these do not appear (see the attached JSON snippet in my original question). There we can see that the spans in indices 1 and 2 have additional key-values (`text`, `source`, and `_input_hash`) while all other spans only have `start`, `end`, `token_start`, `token_end` and `label`.

Why is it that some spans have additional information? Is that arbitrary?

Thanks!!

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [August 10, 2022, 10:23pm UTC](https://support.prodi.gy/t/annotation-json/5746/4 "2022-08-10T22:23:16Z")

</div>

hi @nnn!

Thanks for your follow up!

> [@NNN](#):
>
> 2.) **Unclear** - I still don't understand why for some of the `spans` there are additional keys (`text` , `source` and `_input_hash` ) whereas for others these do not appear (see the attached JSON snippet in my original question). There we can see that the spans in indices 1 and 2 have additional key-values (`text`, `source`, and `_input_hash`) while all other spans only have `start`, `end`, `token_start`, `token_end` and `label`.

Good point! What recipe did you use to create those annotations, specifically the "COUNTRY" and "JOB\_TITLE" spans?

I suspect you used a `correct` recipe (either `ner.correct` or `span.correct`). The highlighted examples (with the extra keys for `text`, `source`, and `input_hash`) is the normal behavior for **model suggested annotations** from a `correct` recipe.

Perhaps the other spans (don't have the extra keys) were created with the same recipe, but **are "manual" annotations** (i.e., you only highlighted) and **weren't model suggested** since the `en_core_web_lg` doesn't have the custom entities (`"COUNTRY"` and `"JOB_TITLE"`).

Said differently, the extra keys (`text`, `source`, and `input_hash`) are created when annotated using model assisted correction.

If you had a trained NER model that had all the entities and were model suggestions, then you would have all the keys/data.

One caveat: it is possible to not have these extra fields for entity types that were in your model (e.g., an `ORG`) because there could be entities that you manually created and weren't model suggestions.

> [@NNN](#):
>
> Is that arbitrary?

It's not required to train so for purposes of training, it's arbitrary.

However, the data does identify the source of the suggestion (e.g., model). Also, having this info distinguishes it as model suggested ("gold" annotations) because it reflects added confidence that the model would select this label (i.e., the data/model are consistent). So in that way, having this data would tell you this data is more important than manual annotations.

Just curious, have you trained a model by updating the original (e.g., `--base-model en_core_web_lg` to update) with both your custom entities and fine-tuned entities from `en_core_web_lg`?

If so, since you're mixing old and new entity types, make sure to account for potential catastrophic forgetting:

> [@New entity model ruins other entities](https://support.prodi.gy/t/new-entity-model-ruins-other-entities/179/2):
>
> Sorry about the late reply! I think what you’re experiencing might be whats often referred to as the “catastrophic forgetting problem”. As your model is learning about the new entity type, it’s “forgetting” what it has previously learned. In your example, this is pretty significant – but it might be because you’ve trained a completely new entity, so the only data the model is updating on is examples labelled TECH and none of the other entity types. Because the model is never “reminded” about the…

I hope this answers your question and let us know if you have other questions!

---

<div class="post-metadata">

**Author:** ![NNN](https://avatars.discourse-cdn.com/v4/letter/n/58f4c7/32.png) [@NNN](https://support.prodi.gy/u/NNN)\
**Post date:** [September 19, 2022, 11:22am UTC](https://support.prodi.gy/t/annotation-json/5746/5 "2022-09-19T11:22:45Z")

</div>

Hi Ryan,

Thanks for your reply.

I used the `ner.correct` recipe and added custom labels.  
Thanks so much for the clear explanation - understood.

We're using `en_core`web\_lg` for the predictions in Prodigy but later converting the output to BIO format to train the model using Flair.

A strange thing though, we getting Prodigy predictions for our custom entities. Is that normal? Could `en_core_web_lg` predict on untrained entities?

Many thanks for everything! 🙏🏼

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [September 26, 2022, 7:52pm UTC](https://support.prodi.gy/t/annotation-json/5746/6 "2022-09-26T19:52:14Z")

</div>

> [@NNN](#):
>
> A strange thing though, we getting Prodigy predictions for our custom entities. Is that normal? Could `en_core_web_lg` predict on untrained entities?

Sorry, I don't fully understand.

Is it that the `en_core_web_lg` `ner` model predicts your custom `ner` labels?

That shouldn't happen. You can view what are the labels in your `ner` component by running:

```python
import spacy
nlp = spacy.load("en_core_web_lg")
nlp.get_pipe('ner').labels
# ('CARDINAL', 'DATE', 'EVENT', 'FAC', 'GPE', 'LANGUAGE', 'LAW', 'LOC', 'MONEY', 'NORP', 'ORDINAL', 'ORG', 'PERCENT', 'PERSON', 'PRODUCT', 'QUANTITY', 'TIME', 'WORK_OF_ART')

```

These are the only labels from this component.

Let me know if I misunderstood your problem.

---

<div class="post-metadata">

**Author:** ![NNN](https://avatars.discourse-cdn.com/v4/letter/n/58f4c7/32.png) [@NNN](https://support.prodi.gy/u/NNN)\
**Post date:** [September 28, 2022, 12:07pm UTC](https://support.prodi.gy/t/annotation-json/5746/7 "2022-09-28T12:07:48Z")

</div>

Hi Ryan,

> [@ryanwesslen](#):
>
> Is it that the `en_core_web_lg` `ner` model predicts your custom `ner` labels?

Yes, or at least so it seems 🙂 (snippet attached)

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/6/600e56d232e111979228c0d5f30eda08c7306ad0.png)

The only build-in labels we're using from `en_core_web_lg` are `PERSON` and `ORG` and `PRODUCT`; the rest are custom.

Thank you

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [September 28, 2022, 7:22pm UTC](https://support.prodi.gy/t/annotation-json/5746/8 "2022-09-28T19:22:00Z")

</div>

Just to make sure I double check, can you run this:

```python
import spacy
nlp = spacy.load("en_core_web_lg")
text = "[Provide example text from similar behavior]"
doc = nlp(text)
spacy.displacy.serve(doc, style="ent")

```

If you're still seeing your custom entities, could you try to [disable other components](https://spacy.io/usage/processing-pipelines#disabling) to keep only `ner`?
