# Does Prodigy support HTML annotation for NER

**URL:** <https://support.prodi.gy/t/does-prodigy-support-html-annotation-for-ner/1370>\
**Category:** Uncategorized\
**Tags:** usage, ner\
**Created:** [April 4, 2019, 12:34am UTC](https://support.prodi.gy/t/does-prodigy-support-html-annotation-for-ner/1370 "2019-04-04T00:34:12Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ichenjia](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ichenjia/32/728_2.png) [@ichenjia](https://support.prodi.gy/u/ichenjia)\
**Post date:** [April 4, 2019, 12:34am UTC](https://support.prodi.gy/t/does-prodigy-support-html-annotation-for-ner/1370/1 "2019-04-04T00:34:12Z")

</div>

I have quite a lot of HTML pages. My goal is to train a model in spacy that can recognize entities such as product name, addresses and monetary amount.

My question is does prodigy support annotation on html pages. not raw html code per se, but annotation on rendered html with a custom list of entities.

Thank you!

Jay

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [April 4, 2019, 9:12am UTC](https://support.prodi.gy/t/does-prodigy-support-html-annotation-for-ner/1370/2 "2019-04-04T09:12:57Z")

</div>

Annotating rendered HTML might sound appealing at first, but there's actually not really an easy answer for how the annotations should be resolved back to the underlying raw text and how to ensure that annotations are consistent. After all, what your model will get to see is the raw text.

I discuss some of these considerations in more detail [on this thread](https://support.prodi.gy/t/using-ner-manual-on-html-input/876):

> [@Using ner.manual on HTML Input](https://support.prodi.gy/t/using-ner-manual-on-html-input/876/2):
>
> If you pass in `"<strong>hello</strong>"` , there’s no clear solution for how this should be handled. How should it be tokenized, and what are you _really_ labelling here? The underlying markup or just the text, and what should the character offsets point to? And how should other markup be handled, e.g. images or complex, nested tags?
> 
> Similarly, if you’re planning on training a model later on, that model will also get to see the raw text, including the markup – so if you are working with raw HTML (like, web dumps or something), you usually always want to see the original raw text that the model will be learning from. Otherwise, the model might be seeing data/markup that you didn’t see during annotation, which is always problematic.

One common solution is to write a function that takes raw HTML, strips out the markup, tokenizes the text and stores each token's character offset into the _original raw HTML_. This way, you can work with raw text without markup, while still being able to resolve the character offsets of your annotations back to the original input.

---

<div class="post-metadata">

**Author:** ![nix411](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/nix411/32/2982_2.png) [@nix411](https://support.prodi.gy/u/nix411)\
**Post date:** [November 30, 2022, 9:34pm UTC](https://support.prodi.gy/t/does-prodigy-support-html-annotation-for-ner/1370/3 "2022-11-30T21:34:29Z")

</div>

> [@ines](#):
>
> One common solution is to write a function that takes raw HTML, strips out the markup, tokenizes the text and stores each token’s character offset into the _original raw HTML_. This way, you can work with raw text without markup, while still being able to resolve the character offsets of your annotations back to the original input.

Hi @ines

I just stumpled upon this and I'm trying to wrap my head around it. But I just can't understand 🙂 Do you mind put a few extra words to it? Or add a few lines of code for illustration?

My use case is for span classification, i.e. subhead classification. I know I have to preprocess the html into text at some point but I like to use the HTML in the tokenization process (improved sentence boundaries e.g.). So the input is really HTML and then I want to mark where the subheads are - simply the offsets/positions of characters.

Alternatively I could just preprocess the HTML to text and start highlighting subheads from there but they are a lot harder to find then. And the classification would then depend on using that exact html to text method

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [December 1, 2022, 7:24pm UTC](https://support.prodi.gy/t/does-prodigy-support-html-annotation-for-ner/1370/4 "2022-12-01T19:24:15Z")

</div>

hi @nix411!

I can't speak for Ines, but have you seen @pmbaumgartner's [`spacy-html-tokenizer`](https://github.com/pmbaumgartner/spacy-html-tokenizer)?

It does the first few steps, but not the character offsets. Perhaps you could look at Peter's code and modify some steps.

I suspect (with my limited knowledge) that there's no general solution that's easy, which is reflected in Ines' point.
