# NER for long string

**URL:** <https://support.prodi.gy/t/ner-for-long-string/6065>\
**Category:** Uncategorized\
**Created:** [November 1, 2022, 1:55pm UTC](https://support.prodi.gy/t/ner-for-long-string/6065 "2022-11-01T13:55:37Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [November 2, 2022, 7:44pm UTC](https://support.prodi.gy/t/ner-for-long-string/6065/2 "2022-11-02T19:44:18Z")

</div>

hi @jiebei!

Thanks for your message!

> [@jiebei](#):
>
> we have annotated the **top level (Primary, Secondary)** with the citations from **3 articles using ner.manual (about 1130 manual citations** in the pattern file now).

I'm a little confused by what you mean "about 1,130 manual citations in the pattern file".

Are these individual examples of citations? Would these not be annotated examples?

How did you obtain these annotations? Did you use the `ner.manual` recipe or some other way?

I'll assume that these 1,130 annotations were created by `ner.manual` for my response below. If I'm not right in assuming that, please let me know.

> [@jiebei](#):
>
> after I generated the pattern file based on the 1130 manual annotations and used it with a new article, nothing can be highlighted as a hint.

I can see the challenge for the whole citation being too long. I would add it's not just because they are long, but also because they can be complex (e.g., a lot of punctuation, numbers, etc.).

Typically, patterns are helpful for starting without any annotations. @koaning has a [great PyData talk](https://youtu.be/nJAmN6gWdK8) where he shows a workflow for this:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/7/71bf70fe95350a652ea5eb027f26d9c00f7c16da.png)

In this case, the matcher (pattern) rules help to provide initial annotations on an unlabeled dataset, which then could be used to train a model.

I suspect that "nothing can be highlighted" because you may have errors in your pattern files. Are you able to test on individual pattern and try to run it through spaCy to confirm it works?

If your 1,130 manual annotations were using `ner.manual`, I think you may benefit from ignoring patterns and build an initial model and then use ["model-in-the-loop" training](https://prodi.gy/docs/named-entity-recognition#manual-model) to improve/add new annotations while improving the model.

## Step 1: create a dedicated evaluation dataset

I would recommend you partition your 1,130 manual annotations into a dedicated training and evaluation dataset. You can see this recent post below where I describe why creating a dedicated evaluation set is a good practice when trying to create experiments to improve your model. That post includes a snippet of code that can take an existing Prodigy dataset (let's say it's named `dataset`), and create two new datasets: `train_dataset` and `eval_dataset`. As that post describes, this is important as you keep your evaluation dataset fixed instead of allowing Prodigy to create a new holdout every time your run `prodigy train`.

> [@How much training data for multiclass/multilabel text classification?](https://support.prodi.gy/t/how-much-training-data-for-multiclass-multilabel-text-classification/6060/4):
>
> If you're using --eval-split, you're not creating a dedicated hold out dataset. It's not based on the --base-model. When you use --eval-split, you're allowing Prodigy to take your dataset, and randomly split it. This is okay early on - but if you run multiple experiments, you'll find that you don't have a fixed / set hold out dataset (that is, each time you run, you get a new 20% holdout). This will likely cause weird results on your train-curve where you run it one time, and the curve is incr…

## Step 2: train an initial `ner` model

I would then recommend training a `ner` model and saving the model. I know that you have multiple hierarchies -- which makes it even more challenging -- but I would recommend starting with your top level first.

When you train this model, like that post recommends, you will need to specify both your training data (let's call `train_dataset`) and your evaluation data (`eval_dataset`):

```python
python -m prodigy train model_folder --ner train_dataset,eval:eval_dataset

```

This will save your model to the `model_folder` folder.

## Step 3: use the `ner.correct` for model predictions, not patterns

Then use the `ner.correct` model without patterns on additional unlabeled data. The `ner.correct` is using your ML model, not the patterns, as the initial labels. You will need to provide the location of your model (`model_folder`).

Once you get your new corrected data, you'll likely want to combine it with your initial training data (`train_dataset`) by using the [`db-merge`](https://prodi.gy/docs/recipes#db-merge) command to create one new "combined" training dataset (initial annotations + newly corrected ones).

## Step 4: Retrain your model

With your new combined dataset, try to retrain your full model.

I hope this helps and let us know if you are able to make any progress!

---

_[View the full topic](https://support.prodi.gy/t/ner-for-long-string/6065)._
