# Expanding NER to include neighbouring tokens

**URL:** <https://support.prodi.gy/t/expanding-ner-to-include-neighbouring-tokens/1178>\
**Category:** Uncategorized\
**Tags:** usage, ner, spacy, finance\
**Created:** [February 2, 2019, 3:20pm UTC](https://support.prodi.gy/t/expanding-ner-to-include-neighbouring-tokens/1178 "2019-02-02T15:20:38Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![nix411](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/nix411/32/2982_2.png) [@nix411](https://support.prodi.gy/u/nix411)\
**Post date:** [February 4, 2019, 8:37pm UTC](https://support.prodi.gy/t/expanding-ner-to-include-neighbouring-tokens/1178/3 "2019-02-04T20:37:24Z")

</div>

Hi @ines

Thanks a lot for the helpful answers you always seem to deliver!

> [@ines](#):
>
> I mean, in theory, you could label some data manually that annotates the spans lik this, update the model with it and hope that it adjusts to your new concept of `MONEY` . The question is whether it’s worth it. If the pre-trained model you’re using was trained with an annotation scheme that considered the currency _not_ part of an entity, all its current weights are based on that policy. It might take a lot of work and data to teach it a very different definition of the entity type `MONEY` .
> 
> **Edit:** Just checked and it seems like the annotation scheme does include the currency by default. It just doesn’t seem to be recognised correctly in this case. I can double-check to see how `MONEY` was annotated in the corpus we’re using for English.

Alright. Good to know. I also read this [post](https://support.prodi.gy/t/advice-on-training-ner-models-with-new-entities/1030/8) where you suggest training a new NER model from scratch. I might take that approach as well - at least test and compare. Is there a way to omit some pretrained labels but keep others?

> [@ines](#):
>
> Another thing to think about: What’s your end goal once you have the money entities? Will you be converting them to some type of structured format like `{'amount': 113000000, 'currency': 'SEK'}` ? If so, it might make more sense to leave the entities the way they are and add custom attributes like `._.currency` to them.

That is exactly what I need to do and I think I will go for the attributes approach indeed.

> [@ines](#):
>
> Btw, I’ currently writing spaCy docs for v2.1 and want to include a section on combining statistical models with rules. Would you mind if I used your example or something very similar for this?

Sure. Let me know if I can be of any help. I have lots of data. My data is a4 pages like of HTML reports but I have chopped them up by parsing the HTML and then added the parsed content to `spacy`. Is it possible to give the whole thing to `spacy` and then do some preprocessing to keep track of the origin of the tokens or should I add that logic as a combination of some custom logic and `spacy`?

> [@ines](#):
>
> Also, totally forgot we had a spaCy code example for entity relation extraction for `MONEY` entities: [see here](https://spacy.io/usage/examples#entity-relations) – looks pretty relevant to you?

It doesn't get more relevant than that!

---

_[View the full topic](https://support.prodi.gy/t/expanding-ner-to-include-neighbouring-tokens/1178)._
