# adding custom attribute to doc, having NER use attribute

**URL:** <https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356>\
**Category:** Uncategorized\
**Tags:** ner, spacy\
**Created:** [March 1, 2018, 5:50pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356 "2018-03-01T17:50:20Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![mhigginslp](https://avatars.discourse-cdn.com/v4/letter/m/839c29/32.png) [@mhigginslp](https://support.prodi.gy/u/mhigginslp)\
**Post date:** [March 1, 2018, 5:50pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/1 "2018-03-01T17:50:20Z")

</div>

I have made a script to auto-reject samples that do not match a pattern from a pattern file when annotating - this works okay for some entities, such as zip-code (I have other 5 digit numbers that aren’t zip-codes). But ultimately it relies on giving the model lots of rejected examples of the entity. For other entities the model will have a much harder time learning to reject a token - if doesn’t match the pattern, for instance if our entity has to match one of a large list.

What would be much better is if our model knew whether or not a given span matches a pattern! To that end I have created a custom component, that adds an attribute signifying if the span matches the pattern.

First Question) Will the NER “know” about my custom attributes? i.e. will it be encoded in the tensor as input?

Second Question) When loading the model using ner.batch-train I am getting the following error:  
KeyError: “Can’t find factory for ‘pattern\_detector’.”  
I read your comment from [[Load error after adding custom textcat model to the pipeline](https://support.prodi.gy/t/load-error-after-adding-custom-textcat-model-to-the-pipeline/122)]  
but I don’t understand where/how to add the factory.

```python

class Pattern_Matcher(object):
    def __init__ (self,nlp, label):
        self.vocab = nlp.vocab
        self.entityname = label
        self.label = nlp.vocab.strings[self.entityname]
        self.matcher = Matcher(nlp.vocab)
        self.name = "pattern_detector"
        self.nlp = nlp
        self.fill_matcher_w_patterns()
        Token.set_extension('is_' + self.entityname , default=False)

    def fill_matcher_w_patterns(self):
        pattern_path = '/data/prodigy/patterns/'+self.entityname+'_pattern.jsonl'
        patterns = []
        with open(str(pattern_path), "r") as f:
            for line in f:
                print (line)
                label =json.loads(line)['label']
                p = json.loads(line)['pattern']
                self.matcher.add(label, None, p)
        print('done adding patterns to matcher')

    def __call__ (self,doc):
        matches = self.matcher(doc)
        spans = []
    
        for _, start, end in matches:
            entity = Span(doc,start, end, label=self.label)
            spans.append(entity)
            for token in entity:
                token._.set('is_' + self.entityname, True)
        for span in spans:
            span.merge()

        return doc

def save_model():
    nlp = spacy.load('en_core_web_lg')

    component = Pattern_Matcher(nlp) # initialise component
 
    nlp.add_pipe(component, first =True)
    nlp.factories["pattern_detector"] = lambda nlp, **cfg: Pattern_Matcher(nlp, label,**cfg)
    print('Pipeline', nlp.pipe_names)
    nlp.to_disk('/data/prodigy/models/zip_me')
'''
```

---

<div class="post-metadata">

**Author:** ![honnibal](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/honnibal/32/35_2.png) [@honnibal](https://support.prodi.gy/u/honnibal)\
**Post date:** [March 1, 2018, 8:00pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/2 "2018-03-01T20:00:44Z")

</div>

> [@mhigginslp](#):
>
> encoded in the tensor as input?

Out-of-the-box, no, it won't. The NER model (and other spaCy models) make use of the `norm`, `prefix`, `suffix` and `shape` [lexical attributes](https://spacy.io/usage/spacy-101#vocab). When you load a model, the vocabulary loads pre-computed values for these features for most common words. For other words, the feature functions in `nlp.vocab.lex_attr_getters` are used to compute the features.

You could try hijacking one of these features to see if it helps. For instance, you could stick your pattern matching onto the shape feature as follows:

```python
from spacy.attrs import SHAPE

get_shape = nlp.vocab.lex_attr_getters[SHAPE]
nlp.vocab.lex_attr_getters[SHAPE] = lambda string: get_shape(string) + matches_pattern(string)
for lex in nlp.vocab:
    # Update any cached values.
    lex.shape_ = nlp.vocab.lex_attr_getters[STRING](lex.orth_)

```

Note that if you do hack the features this way, the pre-trained tagger, parser etc models won't behave properly. If you do find the feature useful, you can subclass the `NamedEntityRecognizer` class and change its `Model()` classmethod, which builds the network. But that will be a bit of effort --- it seems better to check that the feature works first.

---

<div class="post-metadata">

**Author:** ![mhigginslp](https://avatars.discourse-cdn.com/v4/letter/m/839c29/32.png) [@mhigginslp](https://support.prodi.gy/u/mhigginslp)\
**Post date:** [March 2, 2018, 2:31pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/3 "2018-03-02T14:31:30Z")

</div>

> [@mhigginslp](#):
>
> When loading the model using ner.batch-train I am getting the following error:
> 
> KeyError: "Can’t find factory for ‘pattern\_detector’."

Can you please respond to this question as no matter what I do it won't matter if I can't get this bit going.  
I think I just need more explicit instructions - how/where to set Language.factories, and does my class need a to\_disk method?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 2, 2018, 2:48pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/4 "2018-03-02T14:48:55Z")

</div>

> [@mhigginslp](#):
>
> I think I just need more explicit instructions - how/where to set Language.factories

When spaCy loads the model, it checks the `pipeline` and will initialise the individual components by calling `Language.create_pipe`, which looks up the respective factory. You should be able to simply write to the `factories` attribute, which is a dictionary:

```python
Language.factories['pattern_detector'] = lambda nlp, **cfg: PatternMatcher(nlp,**cfg) 

```

The easiest way to ship custom code with your model is to package it as a Python package, using the [`spacy package` command](https://spacy.io/api/cli#package). The model's ` __init__.py` and `load()` method can execute any code, and also include your custom pipeline components or any other arbitrary data. (If your component depends on other libraries, you can even specify those in your model package's requirements).

Btw, when you call `spacy.load` with keyword arguments, all of those will be passed through to your model's `load` method. So if you allow `**cfg` parameters on your custom component, you can load your model like this:

```python
nlp = spacy.load('your_model', label='YOUR_LABEL')

```

You could even add more arguments, like the data path, so you don't have to hard-code any of this into your model.

> [@mhigginslp](#):
>
> does my class need a to\_disk method?

It shouldn't _need_ one – at least, `nlp.to_disk` should only call a pipeline component's `to_disk` method if it exists. But if you do want your component to save data, you can add a simple method that takes the path and saves out data to that location. For example, something like this:

```python
def to_disk(self, path):
    patterns_path = path / 'patterns.json'
    patterns = json.dumps(self.patterns)
    patterns_path.open('w', encoding='utf8').write(patterns)

```

For more details, see the API docs of the [abstract base class `Pipe`](https://spacy.io/api/pipe). It also includes examples of the methods you can implement to make your component trainable (although, I guess this should be less relevant in your case).

---

<div class="post-metadata">

**Author:** ![mhigginslp](https://avatars.discourse-cdn.com/v4/letter/m/839c29/32.png) [@mhigginslp](https://support.prodi.gy/u/mhigginslp)\
**Post date:** [March 5, 2018, 2:59pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/5 "2018-03-05T14:59:28Z")

</div>

No luck so far 🙁  
I successfully made and loaded a custom language model package but I still get the same error:  
KeyError: “Can’t find factory for ‘pattern\_detector’.”

Could you please send me a link to a simple example of how to integrate a custom factory into a spacy model and use it.

Thanks!

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 5, 2018, 3:54pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/6 "2018-03-05T15:54:11Z")

</div>

Thanks for updating and sorry if this has been frustrating. I actually felt inspired by this discussion and built a little example model – see below for the full code 😊

One thing that’s important to keep in mind with this approach is that you do need to **package your model** , install it via pip and load it in from the package, rather than loading it in from a path. Otherwise, your custom code in the model’s ` __init__.py` won’t be executed. (If you’re loading from a path, spaCy will only refer to the model’s `meta.json` and not actually run the package.)

```bash
python -m spacy package /your_model /tmp
cd /tmp/your_model-0.0.0
python setup.py sdist
pip install dist/your_model-0.0.0.tar.gz

```

### Code example

My code assumes that your model data directory contains an `entity_matcher` directory with a `patterns.json` file. In my example, I’m using a JSON file with an object keyed by entity label, e.g. `{"GPE": [...]}`. It also adds a custom `via_patterns` attribute to the spans that lets you see whether an entity was added via the matcher when you use the model in spaCy. This is just a little gimmick for example purposes – so you can leave it out if you don’t need it.

I’ve just used one of the default spaCy models, added `"entity_matcher"` to the pipeline in the `meta.json`, and used the following for the model package’s ` __init__.py`:

```python
# coding: utf8
from __future__ import unicode_literals

from pathlib import Path
from spacy.util import load_model_from_init_py, get_model_meta
from spacy.language import Language
from spacy.tokens import Span
from spacy.matcher import Matcher
import ujson

__version__ = get_model_meta(Path( __file__ ).parent)['version']

def load(**overrides):
    Language.factories['entity_matcher'] = lambda nlp, **cfg: EntityMatcher(nlp,**cfg)
    return load_model_from_init_py( __file__ , **overrides)

class EntityMatcher(object):
    name = 'entity_matcher'

    def __init__ (self, nlp, **cfg):
        Span.set_extension('via_patterns', default=False)
        self.filename = 'patterns.json'
        self.patterns = {}
        self.matcher = Matcher(nlp.vocab)

    def __call__ (self, doc):
        matches = self.matcher(doc)
        spans = []
        for match_id, start, end in matches:
            span = Span(doc, start, end, label=match_id)
            span._.via_patterns = True
            spans.append(span)
        doc.ents = list(doc.ents) + spans
        return doc

    def from_disk(self, path, **cfg):
        patterns_path = path / self.filename
        with patterns_path.open('r', encoding='utf8') as f:
            self.from_bytes(f)
        return self

    def to_disk(self, path):
        patterns = self.to_bytes()
        patterns_path = Path(path) / self.filename
        patterns_path.open('w', encoding='utf8').write(patterns)

    def from_bytes(self, bytes_data):
        self.patterns = ujson.load(bytes_data)
        for label, patterns in self.patterns.items():
            self.matcher.add(label, None, *patterns)
        return self

    def to_bytes(self, **cfg):
        return ujson.dumps(self.patterns, indent=2, ensure_ascii=False)

```

You can also modify the code to take a path to a patterns file, instead of loading the patterns from the model data. This depends on whether you want to ship your patterns with the model, or swap them out. (Shipping your data with the model can be nice if you intend to share it with others – so you can send the `.tar.gz` model to someone else on your team, and they’ll be able to just `pip install` and use it straight away.)

As I mentioned above, the `**cfg` settings are passed down to the component from `spacy.load`, so instead of the `from_disk` and `from_bytes` methods, you can also just get the path or a list of patterns from the config parameters:

```python
patterns_path = cfg.get('patterns_path') # get the path and then read it in
patterns = cfg.get('patterns', []) # get a list of patterns

```

```python
nlp = spacy.load('your_model', patterns_path='patterns.json')
nlp = spacy.load('your_model', patterns=[{'LOWER': 'foo'}])

```

Hope this helps!

---

<div class="post-metadata">

**Author:** ![mhigginslp](https://avatars.discourse-cdn.com/v4/letter/m/839c29/32.png) [@mhigginslp](https://support.prodi.gy/u/mhigginslp)\
**Post date:** [March 5, 2018, 9:05pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/7 "2018-03-05T21:05:39Z")

</div>

Closer! I followed your code/file structure and patterns.json.  
I made a small change from 🙂

```python
self.matcher.add(label, None, *patterns)

```

to:

```python
self.matcher.add(label, None, patterns)

```

I added the custom component first in the pipeline. My base model is en\_core\_web\_lg. The model now loads but when you hand it a string as in:

```python
doc = nlp(u'54354 is a zip')

```

I get the following error:

```python
Segmentation fault (core dumped)

```

When I load disabling the ner everything works as expected.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 5, 2018, 9:25pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/8 "2018-03-05T21:25:29Z")

</div>

Ah okay – I think my patterns file included a list of separate patterns for the same label, so I made the component add them all to one match pattern. But you can obviously solve this however you like.

A segfault should never happen, so there’s at least an unhandled error somewhere. One likely explanation could be that something goes wrong when the entity recognizer encounters the already set entities. This _should_ work – but there might be certain edge cases where it fails. (If I remember correctly, there’s currently an open issue about a similar problem on the spaCy tracker, which we weren’t able to reproduce. But your example here is actually very nice and isolated – so if you’re able to share an example, we might be able to get to the bottom of this bug 💪)

One thing you could try in the meantime is to add the component _last_, i.e. after the entity recognizer. Before modifying the `doc.ents` in your component’s ` __call__ ` method, you could also check whether it already includes an entity for that exact span (or an overlapping span) and filter that out. In your case, this should be fairly easy – I’m pretty sure your zip codes will be recognised as separate `ORDINAL` entities.

---

<div class="post-metadata">

**Author:** ![mhigginslp](https://avatars.discourse-cdn.com/v4/letter/m/839c29/32.png) [@mhigginslp](https://support.prodi.gy/u/mhigginslp)\
**Post date:** [March 6, 2018, 5:16pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/9 "2018-03-06T17:16:22Z")

</div>

I don’t think there is anything to share - all the code you provided! The pattern file is just:

```python
{ "ZIPCODE": [{"IS_DIGIT":true ,"LENGTH":5}] }

```

I also tried the same thing with another pattern just seeing if the text is a particular string and the same behavior occured.

When I feed in text that does not match the pattern as in:

```python
doc = nlp('This text does not contain a zipcode')

```

the segmentation error does not occur.

Another thing:  
When I run ner.batch-train with the label ZIPCODE I do not get the error … the model trains per usual!  
Unfortunately as Honnibal pointed out the NER does not use custom attributes so there is no improved performance, so as suggested I will try hijacking the shape\_ attribute of the vocab.

---

<div class="post-metadata">

**Author:** ![mhigginslp](https://avatars.discourse-cdn.com/v4/letter/m/839c29/32.png) [@mhigginslp](https://support.prodi.gy/u/mhigginslp)\
**Post date:** [March 6, 2018, 10:30pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/10 "2018-03-06T22:30:37Z")

</div>

> [@honnibal](#):
>
> You could try hijacking one of these features to see if it helps. For instance, you could stick your pattern matching onto the shape feature as follows:
> 
> from spacy.attrs import SHAPE
> 
> get\_shape = nlp.vocab.lex\_attr\_getters[SHAPE]  
> nlp.vocab.lex\_attr\_getters[SHAPE] = lambda string: get\_shape(string) + matches\_pattern(string)  
> for lex in nlp.vocab:  
> # Update any cached values.  
> lex.shape\_ = nlp.vocab.lex\_attr\_gettersSTRING

Unfortunately, this causes a stack overflow error ---- matches\_pattern(string) must use a Matcher object which uses the lex\_attr\_getter function for the shape\_ of the input.  
Another issue is that changing the shape of the vocab will only be effective if the terms I want to change happen to be in the vocabury -- in the case of 5 number digits this is going to be infrequent.

I ended up using:

```python
for i in range(100000):
     new_zip = u"{0:0=5d}".format(i)
     nlp.vocab.set_vector(new_zip, np.zeros(300) )
     nlp.vocab[new_zip].shape_ = u'IMAZIPCODE'

```

You could do something similar if you are looking to only match terms from your patterns file.

---

<div class="post-metadata">

**Author:** ![honnibal](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/honnibal/32/35_2.png) [@honnibal](https://support.prodi.gy/u/honnibal)\
**Post date:** [March 9, 2018, 5:16pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/11 "2018-03-09T17:16:21Z")

</div>

> [@mhigginslp](#):
>
> Another issue is that changing the shape of the vocab will only be effective if the terms I want to change happen to be in the vocabury – in the case of 5 number digits this is going to be infrequent.

The `lex_attr_getter` functions are called on each new word, if it's not found in the vocab. So if you set your shape-setting function in `vocab.lex_attr_getters[SHAPE]` it should work.

I'm really confused by this segfault. I can't see what would be going on --- the following works fine for me:

```python
>>> from spacy.matcher import Matcher
>>> import spacy
>>> nlp = spacy.load('en_core_web_sm')
>>> doc = nlp(u'90210 is a zip code')
>>> m = Matcher(nlp.vocab)
>>> m.add('ZIPCODE', None, [{"IS_DIGIT":True ,"LENGTH":5}])
>>> m(doc)
[(12564065238629231850, 0, 1)]

```

So there must be something else going on. I normally debug segfaults by commenting out parts of the program until it runs successfully, and then commenting back in. It's sort of a brutal approach, but usually the binary search only takes a few tries.

So does it run if you don't do this bit?

> [@mhigginslp](#):
>
> ```python
> for _, start, end in matches:
> entity = Span(doc,start, end, label=self.label)
> spans.append(entity)
> for token in entity:
> token._.set('is_' + self.entityname, True)
> for span in spans:
> span.merge()
> 
> ```

If that runs, I think the problem is in the merging. It's hard to make sure that the token offsets stay valid after merging spans, especially if we have some spans overlapping. Try this:

```python
def __call__ (self,doc):
    matches = self.matcher(doc)
    spans = []
    for _, start, end in matches:
        entity = Span(doc,start, end, label=self.label)
        spans.append((entity.start_char, entity.end_char))
        for token in entity:
            token._.set('is_' + self.entityname, True)
    for start_char, end_char in spans:
        doc.merge(start_char, end_char, ent_type=self.label)
    return doc

```

---

<div class="post-metadata">

**Author:** ![honnibal](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/honnibal/32/35_2.png) [@honnibal](https://support.prodi.gy/u/honnibal)\
**Post date:** [March 9, 2018, 5:19pm UTC](https://support.prodi.gy/t/adding-custom-attribute-to-doc-having-ner-use-attribute/356/12 "2018-03-09T17:19:45Z")

</div>

> [@ines](#):
>
> One thing you could try in the meantime is to add the component last, i.e. after the entity recognizer. Before modifying the doc.ents in your component’s **call** method, you could also check whether it already includes an entity for that exact span (or an overlapping span) and filter that out. In your case, this should be fairly easy – I’m pretty sure your zip codes will be recognised as separate ORDINAL entities.

This is the other possible explanation: instead of the merging, we might be getting an error from the NER running over the text.

Actually, a though. I wonder whether the `.merge()` method is failing to set the `.ent_iob` attribute correctly when we supply an entity type during merging. So maybe it's the combination. I'll check.

Update: Yes that seems to be the case. 🎉  
Happy to have gotten to the bottom of this --- it's been a problem for spaCy for a while. It's hard to debug because it's the combination merge + set entity type + apply NER + conflicting prediction.
