# ner.manual pattern file

**URL:** <https://support.prodi.gy/t/ner-manual-pattern-file/4595>\
**Category:** Uncategorized\
**Tags:** usage, ner\
**Created:** [August 20, 2021, 5:49am UTC](https://support.prodi.gy/t/ner-manual-pattern-file/4595 "2021-08-20T05:49:25Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dsr2021](https://avatars.discourse-cdn.com/v4/letter/d/dec6dc/32.png) [@dsr2021](https://support.prodi.gy/u/dsr2021)\
**Post date:** [August 20, 2021, 5:49am UTC](https://support.prodi.gy/t/ner-manual-pattern-file/4595/1 "2021-08-20T05:49:25Z")

</div>

I am using a regex pattern file to assist prodigy in finding patterns for MONEY using regex using ner.manual and en\_core\_web\_lg and a labels file that includes DATE, LOC, GPE, and more. Prodigy correctly identifies all the patterns that I have set for MONEY. It doesn't seem to annotate anything else! Does adding a pattern file stop prodigy from identifying DATE, LOC, GPE? I was under the impression that a patterns file was a a helper to improve prediction and not a replacement for a complete set of recognizing rules. What do I need to do to allow the usual recognition of entities, and create patterns to improve recognition?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [August 21, 2021, 3:02am UTC](https://support.prodi.gy/t/ner-manual-pattern-file/4595/2 "2021-08-21T03:02:06Z")

</div>

Hi! The `ner.manual` workflow is fully manual and won't use the model for its predictions (only for tokenization). The main use case of patterns here is to help you pre-label common instances so you don't have to do _everything_ full from scratch.

The `ner.correct` workflow will stream in the model's predictions and will let you correct them manually. However, it doesn't have an option for patterns because that'd introduce a slightly tricky question about how to deal with conflicts and overlaps, which can often happen. I explain this in more detail in this thread:

> [@spans.correct recipe](https://support.prodi.gy/t/spans-correct-recipe/4562/2):
>
> This is a bit more difficult and introduces the problem of how matches vs. predictions should be handled, and which to prefer in case there are overlaps. For NER uses cases, you could default to showing either the prediction or pattern match if they disagree – although, it's often useful to see both, but you still want to make sure that your final data ends up with only one version. And while the span categorizer can predict overlapping spans, you'd often still want to pick one span that's most consistent. For instance, the model may predict a "the" + noun phrase, while your pattern describes only the noun phrase. In that case, you want to make sure that your final data only ends up with one of them, not both. The "comparing annotations" workflow described in this issue goes in a similar direction, and it's definitely something you could implement in a custom recipe: [Recipe for comparing NER model and manual annotation - #3 by haishao](https://support.prodi.gy/t/recipe-for-comparing-ner-model-and-manual-annotation/4425/3)

One option could be to use `ner.manual` with patterns for your new categories, and `ner.correct` with the existing model for all others. When you train your model, Prodigy will automatically merge all annotations on the same text, so it's fine if you have the same example annotated twice with different labels.

Alternatively, you could also implement a small variation of [`ner.manual`](https://github.com/explosion/prodigy-recipes/blob/master/ner/ner_manual.py) that also includes the predictions – you just need to make sure that the data you send out doesn't include any overlaps. You could just filter the spans using spaCy's [`filter_spans`](https://spacy.io/api/top-level#util.filter_spans) utility and prefer whatever comes first if there's a conflict. Alternatively, you could also decide to prefer pattern matches over predictions, or vice versa. Or you could decide this on a per-label basis – ultimately, this depends on the data and what types of conflicts are most common.
