# terms.to-patterns looks strange

**URL:** <https://support.prodi.gy/t/terms-to-patterns-looks-strange/916>\
**Category:** Uncategorized\
**Tags:** terms, solved\
**Created:** [October 23, 2018, 3:33pm UTC](https://support.prodi.gy/t/terms-to-patterns-looks-strange/916 "2018-10-23T15:33:45Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mindaugas](https://avatars.discourse-cdn.com/v4/letter/m/48db29/32.png) [@Mindaugas](https://support.prodi.gy/u/Mindaugas)\
**Post date:** [October 23, 2018, 3:33pm UTC](https://support.prodi.gy/t/terms-to-patterns-looks-strange/916/1 "2018-10-23T15:33:45Z")

</div>

Hi guys,

I am very new to Prodigy and while exploring a suspicious thing got my attention. I tried to improve the existing model and followed the steps described here ([https://prodi.gy/docs/workflow-first-steps](https://prodi.gy/docs/workflow-first-steps)). So basically:  
python -m prodigy dataset my\_set “Playground” --author Me  
python -m prodigy ner.teach my\_set en\_core\_web\_sm news\_headlines.jsonl --label ORG  
python -m prodigy terms.to-patterns my\_set out.jsonl --label ORG

Exported patterns look like:  
{“label”:“ORG”,“pattern”:[{“lower”:“The War Between Apple and Google Has Just Begun”}]}  
{“label”:“ORG”,“pattern”:[{“lower”:“Uber\u2019s Lesson: Silicon Valley\u2019s Start-Up Machine Needs Fixing”}]}

My question is are the patterns correct? I would expect patterns to be something like (as I accepted only names of companies as an entity of organization):  
{“label”:“ORG”,“pattern”:[{“lower”:“Apple”}]}  
{“label”:“ORG”,“pattern”:[{“lower”:“Google”}]}  
Thanks in advance.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [October 23, 2018, 3:57pm UTC](https://support.prodi.gy/t/terms-to-patterns-looks-strange/916/2 "2018-10-23T15:57:15Z")

</div>

Sorry if this was confusing – the `terms.to-patterns` recipe is designed to convert a dataset of _single terms_ to a patterns file – fore example, a dataset created with `terms.teach`, which would include examples like `"text": "Apple"`. That patterns file can then be used to bootstrap training in `ner.teach` and make sure the model sees enough positive suggestions.

Creating patterns from existing annotations is a good idea, though – you could even use `ner.manual` to label a few texts manually and then convert the highlighted spans to patterns. There's no built-in recipe for this, but writing your own converter is pretty straightforward. Essentially, all you have to do is load the dataset, get the accepted annotations and use the `"spans"` property (highlighted text) to extract the entity text and add it to the list of patterns:

```python
from prodigy.components.db import connect
from prodigy.util import write_jsonl

db = connect() # connect to DB with setting sfrom prodigy.json
examples = db.get_dataset('my_set') # load the dataset

patterns = []
for eg in examples: # iterate over the annotations
    if eg['answer'] == 'accept': # we only want accepted entities
        spans = eg.get('spans', []) # get the annotated spans
        for span in spans:
            # get the highlighted text and create a pattern
            text = eg['text'][span['start']:span['end']]
            patterns.append({'pattern': text, 'label': span['label']})

write_jsonl('/path/to/patterns.jsonl', patterns)

```

The above example only creates patterns for exact string matches, e.g. `"pattern": "Apple"`. If you want case-insensitive token-based matching, you can use spaCy to tokenize the text for you and create a pattern this way:

```python
text = eg['text'][span['start']:span['end']]
doc = nlp(text)
tokens = [{'lower': token.lower_} for token in doc]
patterns.append({'pattern': tokens, 'label': span['label']})

```

You can also check out this thread, which discusses a similar approach and solution for creating patterns:

> [@Train a new NER entity with multi-word tokens](https://support.prodi.gy/t/train-a-new-ner-entity-with-multi-word-tokens/227):
>
> I’m not sure if I’m dealing with a bug or if I’m doing something wrong. But also if it’s the latter case you might be interested in why I’m doing this, so I’m going to write a quite verbose message here that will let you follow my thought process. My goal is to train a new NER entity with the name DISASTER which will recognize for example floods, storms and volcano eruptions. Yesterday I followed your video in which you train a DRUG entity and got the results I wanted in the end. But now I’m st…

---

<div class="post-metadata">

**Author:** ![Mindaugas](https://avatars.discourse-cdn.com/v4/letter/m/48db29/32.png) [@Mindaugas](https://support.prodi.gy/u/Mindaugas)\
**Post date:** [October 23, 2018, 7:19pm UTC](https://support.prodi.gy/t/terms-to-patterns-looks-strange/916/3 "2018-10-23T19:19:14Z")

</div>

Things are starting to make sense. Thanks for a quick answer!
