# Store the annotation obtained by ner.manual and --patterns at once

**URL:** <https://support.prodi.gy/t/store-the-annotation-obtained-by-ner-manual-and-patterns-at-once/4350>\
**Category:** Uncategorized\
**Tags:** usage, ner, spacy, solved\
**Created:** [June 22, 2021, 11:52am UTC](https://support.prodi.gy/t/store-the-annotation-obtained-by-ner-manual-and-patterns-at-once/4350 "2021-06-22T11:52:31Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![fsa](https://avatars.discourse-cdn.com/v4/letter/f/e9a140/32.png) [@fsa](https://support.prodi.gy/u/fsa)\
**Post date:** [June 22, 2021, 11:52am UTC](https://support.prodi.gy/t/store-the-annotation-obtained-by-ner-manual-and-patterns-at-once/4350/1 "2021-06-22T11:52:31Z")

</div>

I am using the ner.manual recipe with patterns to annotate a given text

> `prodigy ner.manual dataset spacy_model source --label --patterns`

However, currently I don't like to go through all annotations and confirm them one by one instead I want to store at once in the given database all matched entities/labels with the one in the patterns file. The initial annotation I will then use to build the model in an active learning scenario.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 23, 2021, 2:00am UTC](https://support.prodi.gy/t/store-the-annotation-obtained-by-ner-manual-and-patterns-at-once/4350/2 "2021-06-23T02:00:47Z")

</div>

Hi! In that case, you could just load the patterns with spaCy directly to label all matches automatically and then use that data to pretrain you model. My comment here explains how to do this:

> [@How to perform automatically NER annotation based on patterns?](https://support.prodi.gy/t/how-to-perform-automatically-ner-annotation-based-on-patterns/4284/2):
>
> Hi! In that case, you can just go directly via spaCy, for example, using the entity ruler: [https://spacy.io/usage/rule-based-matching#entityruler](https://spacy.io/usage/rule-based-matching#entityruler) It lets you add your patterns and will add all matches to the doc.ents, just like an entity recognizer. You can then use that nlp object to process your texts and extract the pattern-based NER annotations. In theory, you don't even have to go through Prodigy at all and you could just export the data and train with spaCy directly. But if you want to mi…

Using the `EntityRuler` has the advantage that it takes patterns in the same format as Prodigy and takes care of filtering out overlaps (which can theoretically occur with multiple patterns).

---

<div class="post-metadata">

**Author:** ![fsa](https://avatars.discourse-cdn.com/v4/letter/f/e9a140/32.png) [@fsa](https://support.prodi.gy/u/fsa)\
**Post date:** [June 25, 2021, 9:20am UTC](https://support.prodi.gy/t/store-the-annotation-obtained-by-ner-manual-and-patterns-at-once/4350/3 "2021-06-25T09:20:35Z")

</div>

@ines Thanks a lot !  
it works now

Here I share my experience:  
My source data is in jsonl format and look like:

> {"text":"abcd","meta":{"source":"doc1"}}  
> .  
> .

I wrote a code (compatible with SpaCy 2.5) based on your explanation to read a set of documents and annotate them based on patterns file:

```python
# path of jsonl file contains the performed annotation to be loaded in the db
db_jsonl_path='db_jsonl.jsonl'
nlp = English()
ruler = EntityRuler(nlp)
# the patterns file
ruler.from_disk('patterns.jsonl') 
nlp.add_pipe(ruler)

# source data in jsonl format
source_path='soure_data.jsonl'
# Using readlines()
source_file = open(source_path, 'r')
Lines = source_file.readlines()
 
for line in Lines:
    data = json.loads(line.strip())
    input=data['text']
    doc = nlp(input)
    spans = [{"start": ent.start_char, "end": ent.end_char, "label": ent.label_} for ent in doc.ents]
    example = {"text": doc.text, "spans": spans}
    with open(db_jsonl_path, 'w') as f:
                       f.write(json.dumps(example+'\n')

```

when done, load the performed annotation, stored in db\_jsonl\_path, into a prodigy db:

```python
prodigy db-in db_name path/db_jsonl.jsonl

```

I still have a simple question, how to add the meta data ("meta":{"source":"doc1"}) into the spans so it can be stored in the db later a long with other information like entities, position, label etc.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 28, 2021, 2:02am UTC](https://support.prodi.gy/t/store-the-annotation-obtained-by-ner-manual-and-patterns-at-once/4350/4 "2021-06-28T02:02:50Z")

</div>

> [@fsa](#):
>
> I still have a simple question, how to add the meta data ("meta":{"source":"doc1"}) into the spans so it can be stored in the db later a long with other information like entities, position, label etc.

You can add all of that to the dict that you create as `example` in your code 🙂 The `"text"` and `"spans"` are what's required to annotate named entities, but you can also include a key `"meta"` with custom properties – for example, the index of the current line (you can just increment a counter variable or use Python's `enumerate()`).

Everything in `"meta"` will be be displayed in the bottom right corner of the annotation card. You can also include any other custom properties in the example that will be saved with the annotations in the database (e.g. for meta infor that you don't want to display in the UI).

---

<div class="post-metadata">

**Author:** ![fsa](https://avatars.discourse-cdn.com/v4/letter/f/e9a140/32.png) [@fsa](https://support.prodi.gy/u/fsa)\
**Post date:** [June 28, 2021, 9:12am UTC](https://support.prodi.gy/t/store-the-annotation-obtained-by-ner-manual-and-patterns-at-once/4350/5 "2021-06-28T09:12:19Z")

</div>

Thanks a lot, I added the meta data and it is displayed in the bottom right corner of the annotation card:

```python
# path of jsonl file contains the performed annotation to be loaded in the db
db_jsonl_path='db_jsonl.jsonl'
nlp = English()
ruler = EntityRuler(nlp)
# the patterns file
ruler.from_disk('patterns.jsonl') 
nlp.add_pipe(ruler)

# source data in jsonl format
source_path='soure_data.jsonl'
# Using readlines()
source_file = open(source_path, 'r')
Lines = source_file.readlines()
 
for line in Lines:
    data = json.loads(line.strip())
    input=data['text']
    doc_id=data['meta']
    doc = nlp(input)
    spans = [{"start": ent.start_char, "end": ent.end_char, "label": ent.label_} for ent in doc.ents]
    example = {"text": doc.text, "spans": spans,"meta":doc_id}
    with open(db_jsonl_path, 'a') as f:
                       f.write(json.dumps(example+'\n')

```
