# Pre-annotate entities with patterns

**URL:** <https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394>\
**Category:** Uncategorized\
**Tags:** usage, ner, solved\
**Created:** [July 2, 2021, 8:07am UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394 "2021-07-02T08:07:56Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![alvaro.marlo](https://avatars.discourse-cdn.com/v4/letter/a/b38774/32.png) [@alvaro.marlo](https://support.prodi.gy/u/alvaro.marlo)\
**Post date:** [July 2, 2021, 8:07am UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394/1 "2021-07-02T08:07:57Z")

</div>

Hello,

I want to create an annotations jsonl file using patterns with es\_core\_news\_lg pipeline. In this case, I'm trying to annotate MONEY entities using a jsonl file I have created with the patterns.

There's any possibility to do that? Probably using custom recipes? I tried using get\_stream for a data jsonl file and PatternMatcher.from\_disk for the patterns jsonl file (there is the code below). But the problem is that if the text has more than one entity, it creates different json's in the jsonl preannotation file. How can I solve the problem?

```python
import spacy
import prodigy
from prodigy.components import printers
from prodigy.components.loaders import get_stream
from prodigy.core import recipe, recipe_args
from prodigy.models.matcher import PatternMatcher
from prodigy.util import log
import json

@prodigy.recipe('ner.preannotate-patterns',
        spacy_model=recipe_args['spacy_model'],
        patterns=('Path to match patterns file', 'positional'),
        source=recipe_args['source'],
        api=recipe_args['api'],
        loader=recipe_args['loader'])
def print_pattern_stream(spacy_model, patterns, source=None, api=None, loader=None):
    log("RECIPE: Starting recipe ner.preannotate-patterns", locals())
    model = PatternMatcher(spacy.load(spacy_model)).from_disk(patterns)
    stream = get_stream(source, api, loader, rehash=True, input_key='text')
    with open('preannotations_money.jsonl', 'w') as file:
        for line in model(stream):
            json_line = json.dumps(line[1])
            file.write(json_line)
            file.write("\n")
        file.close()

```

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [July 3, 2021, 2:47am UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394/2 "2021-07-03T02:47:22Z")

</div>

Hi! From what you describe, it sounds like you could solve this by using `ner.manual` with `--patterns`: [Named Entity Recognition · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/named-entity-recognition#manual-patterns)

Prodigy supports both token-based and extract string patterns (like the ones used by spaCy's `PatternMatchers`), so your file could look like this:

```json
{"label": "MONEY", "pattern": "€123"}
{"label": "MONEY", "pattern": [{"IS_CURRENCY": true}, {"LIKE_NUM": true}]}

```

> [@alvaro.marlo](#):
>
> But the problem is that if the text has more than one entity, it creates different json's in the jsonl preannotation file.

Just a quick note on how you would solve this if you wanted to implement this in a custom recipe – although you shouldn't have to because `ner.manual` has you covered 🙂

Under the hood, the pre-annotation works like this:

- Stream in all your input examples and load the patterns into a matcher.
- For each example in the stream, process it with spaCy and match your patterns on it.
- If relevant: Make sure that you filter the matches so you don't end up with overlapping spans. You can do this using spaCy's [`filter_spans`](https://spacy.io/api/top-level#util.filter_spans) utility.
- Add a list of `"spans"` to each example containing the matches. spaCy's matcher will let you create `Span` objects for each match, so you can do `{"start": span.start_char, "end": span.end_char, "label": "MONEY"}` for each span.
- Add tokens your stream using Prodigy's [`add_tokens`](https://prodi.gy/docs/api-components#add_tokens) helper.

---

<div class="post-metadata">

**Author:** ![ben.k](https://avatars.discourse-cdn.com/v4/letter/b/e56c9b/32.png) [@ben.k](https://support.prodi.gy/u/ben.k)\
**Post date:** [January 5, 2023, 12:21pm UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394/3 "2023-01-05T12:21:21Z")

</div>

I am using `ner.manual` with `--patterns` but I wanted to add to a larger JSONL-file (which includes many patterns already) another pattern that would pre-annotate money amounts as predicted by the `en_core_web_sm` model. When using this pattern `{"label": "MONEY", "pattern": [{"ent_type": "MONEY"}]}`, however, I am getting this result below:

![Screenshot from 2023-01-05 13-10-45](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/f/fa97ccd5642d6880e69ec521e8f3ac47909b0c3c.png)

What I had in mind was that `$ 75 million` is pre-annotated as `MONEY` combined (instead of the individual elements) like the `en_core_web_sm` model would typically do in the case of NER.

To clarify, I am not using `ner.correct` as shown under [this link](https://prodi.gy/docs/named-entity-recognition#manual-model) since it does not allow me to add a patterns file in which I would collect additional patterns than the MONEY entity described above.

How can I combine both, the classical patterns collected in a JSONL-file and named entities (such as `MONEY`) prediced by an existing model, for the task of pre-annotating a text?

Many thanks! Ben

---

<div class="post-metadata">

**Author:** ![Jette16](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/jette16/32/3162_2.png) [@Jette16](https://support.prodi.gy/u/Jette16)\
**Post date:** [January 10, 2023, 8:24am UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394/4 "2023-01-10T08:24:24Z")

</div>

Hello @ben.k,  
thank you for your question and welcome to the prodigy community 👋

The problem with the MONEY label results from your pattern and the model's tokenization. As decribed in [the prodigy docs](https://prodi.gy/docs/api-loaders#input-patterns), a pattern is a list of dictionaries where each dictionary describes one individual token. Your pattern `{"label": "MONEY", "pattern": [{"ent_type": "MONEY"}]}` is just one token long, but because spaCy tokenizes `$ 75 million` into three tokens, each having the entity type `MONEY`, your pattern is matched three times leading to the result you described.

One solution to solve this could be using the `"OP"` key in your pattern, see [https://spacy.io/usage/rule-based-matching#quantifiers](https://spacy.io/usage/rule-based-matching#quantifiers). For example,

```python
{"label": "MONEY", "pattern": [{"ent_type": "MONEY", "OP": "+"}]}

```

as well as

```python
{"label": "MONEY", "pattern": [{"ent_type": "MONEY"},{"ent_type": "MONEY", "op": "?"},{"ent_type": "MONEY", "op": "?"}]}

```

result in a combined MONEY label.

Another solution could be to write a custom recipe to solve this. However, since your problem should be solvable using the operators, I would not recommend this.

I hope this solves your problem. If not or if you have any further questions, please let me know 🙂

---

<div class="post-metadata">

**Author:** ![ben.k](https://avatars.discourse-cdn.com/v4/letter/b/e56c9b/32.png) [@ben.k](https://support.prodi.gy/u/ben.k)\
**Post date:** [January 10, 2023, 3:37pm UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394/5 "2023-01-10T15:37:17Z")

</div>

Hi @Jette16 ,

many thanks for your answer, this is exactly what I was looking for. Now I understand the role of the `"OP"` key, it is similar to the procedure in classical REGEX patterns, and it works as expected.

One addition question. Now that spaCy has recognised `$ 75 million` as a named entity of type money, is there (by any chance) a utility function that allows me to split this into the currency unit and the number, i.e. `('USD', 75000000)`, or would I do that using a procedure as described e.g. [here](https://stackoverflow.com/questions/48359255/python-match-and-parse-strings-containing-numeric-currency-amounts)?

In principle, with my initial pattern the different elements were labelled separately, but I would need to know whether each element is a currency classifier, the amount or the unit multiplier to proceed further. Hence, maybe it is easier to combine it using the `"OP"` key and then pass the entire recognised named entity into such parser-function.

---

<div class="post-metadata">

**Author:** ![Jette16](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/jette16/32/3162_2.png) [@Jette16](https://support.prodi.gy/u/Jette16)\
**Post date:** [January 11, 2023, 1:59pm UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394/6 "2023-01-11T13:59:11Z")

</div>

Hi @ben.k,  
I'm happy to hear that the proposed solution works 🙂

Regarding your question:  
If you have a solution that already works, I would go with it. Prodigy does not have such a utility function.  
What you could do is to incorporate this parser function into your own prodigy recipe. Another option could be to extend the patterns you already wrote, like adding a regex key and split the pattern for MONEY up into two parts like this:

```python
{"label": "USD", "pattern": [{"ent_type": "MONEY", "text": {"REGEX": "[$]"}}]}
{"label": "AMOUNT", "pattern": [{"ent_type": "MONEY", "text": {"REGEX": "[^$]"}},{"ent_type": "MONEY", "op": "?", "text": {"REGEX": "[^$]"}}]}

```

However, depending on your text and the variations that you'd have to cover (e.g., different currencies,...), this might be an overhead, especially since you'd still have to convert "75 million" to "75000000".

A third possibility could be to create your own [spaCy pipeline component](https://spacy.io/usage/processing-pipelines#custom-components) to implement this step which might be useful if you already have a processing step that utilizes spaCy. If you have more specific questions to custom components, I'd like to refer to the [spaCy Discussions forum](https://github.com/explosion/spaCy/discussions) where my colleagues will help you with your spaCy questions.

I hope one of the possibilities suits you. Like I said, I would use the parser that you proposed, especially if you have already implemented this.

---

<div class="post-metadata">

**Author:** ![ben.k](https://avatars.discourse-cdn.com/v4/letter/b/e56c9b/32.png) [@ben.k](https://support.prodi.gy/u/ben.k)\
**Post date:** [January 11, 2023, 9:52pm UTC](https://support.prodi.gy/t/pre-annotate-entities-with-patterns/4394/7 "2023-01-11T21:52:49Z")

</div>

Many thanks for your comprehensive reply, including the examples on how to split currency unit and amount using spaCy!
