# terms.teach hangs indefinitely with a custom word vector model

**URL:** <https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106>\
**Category:** Uncategorized\
**Tags:** terms, solved\
**Created:** [January 5, 2019, 5:12pm UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106 "2019-01-05T17:12:49Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![bigbeaker](https://avatars.discourse-cdn.com/v4/letter/b/7feea3/32.png) [@bigbeaker](https://support.prodi.gy/u/bigbeaker)\
**Post date:** [January 5, 2019, 5:12pm UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/1 "2019-01-05T17:12:49Z")

</div>

I have my custom spacy model uses custom word vectors.

The word vectors work fine in the spacy model:

```python
nlp = spacy.load('test_model')
nlp.vocab.length
467868
nlp.vocab.vectors_length
300
nlp.vocab.has_vector('وحش')
True

```

Loading the same model with prodigy:

```python
pgy terms.teach dataset test_model -s "وحش"
Initialising with 1 seed terms: وحش

  ✨ Starting the web server at http://localhost:8080 ...
  Open the app in your browser and start annotating!

```

All looks fine, no errors, but I don’t get any terms to annotate, it just hangs with a …Loading message on the webpage.

Any insight or further checks I can do would be greatly appreciated.

All data is in Arabic, but that doesn’t seem to be an issue with other parts of prodigy!

Spacy version: 2.0.18  
Prodigy version: 1.6.1

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 5, 2019, 5:57pm UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/2 "2019-01-05T17:57:25Z")

</div>

Hi! I think I know what might be going on here: When `terms.teach` loops over the vocabulary in the vocab, it does the following:

```python
lexemes = [lex for lex in stream if lex.is_alpha and lex.is_lower]

```

`is_lower` actually returns `False` for the arabic tokens – it delegates to Python’s native `islower()`, which is also `False`. (Interestingly, Python’s `isupper()` is `False`, too. I guess all of this is logical, because there’s no uppercase and lowercase distinction in Arabic, right?)

So as a quick fix, removing `and lex.is_lower` in `recipes/terms.py` should do the trick.

---

<div class="post-metadata">

**Author:** ![bigbeaker](https://avatars.discourse-cdn.com/v4/letter/b/7feea3/32.png) [@bigbeaker](https://support.prodi.gy/u/bigbeaker)\
**Post date:** [January 5, 2019, 7:17pm UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/3 "2019-01-05T19:17:02Z")

</div>

Hey Ines

> [@ines](#):
>
> I guess all of this is logical, because there’s no uppercase and lowercase distinction in Arabic, right?)

Yes correct, there's no upper/lower case in Arabic.  
is\_alpha returns True

Changed to:

```python
lexemes = [lex for lex in stream if lex.is_alpha]

```

and it worked!! 💪💪  
Thanks a lot Ines, appreciate the quick reply.

Side note: the prodigy package is written beautifully.

---

<div class="post-metadata">

**Author:** ![bigbeaker](https://avatars.discourse-cdn.com/v4/letter/b/7feea3/32.png) [@bigbeaker](https://support.prodi.gy/u/bigbeaker)\
**Post date:** [January 6, 2019, 1:31am UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/4 "2019-01-06T01:31:10Z")

</div>

Follow up question, does this also affect the patterns matching functionality:

```python
{"label":"NEGATIVE","pattern":[{"lower":"قرف"}]}

```

As I don’t seem to be getting any pattern matches with:

```python
prodigy textcat.teach classifier spacy_model data.txt --label NEGATIVE --patterns seed_file.jsonl

```

Checked with spacy & I do get matches with the LOWER pattern on Arabic:

```python
matcher = Matcher(nlp.vocab)
pattern = [{'LOWER': "قرف"}]
matcher.add("HelloWorld", None, pattern)

doc = nlp('خر قرف شي طايس')
matches = matcher(doc)
for match in matches:
    print(match)

```

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 7, 2019, 10:12am UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/5 "2019-01-07T10:12:08Z")

</div>

> [@bigbeaker](#):
>
> Follow up question, does this also affect the patterns matching functionality:

In theory, it shouldn't – matching on the `lower` attribute just means that the matcher will compare the lowercase forms of both tokens (which should be identical either way – I just checked for the string you provided and it seems like `token.lower_ == token.text` in the case of Arabic).

Under the hood, Prodigy calls into spaCy's `Matcher` and `PhraseMatcher`. Can you double-check you're on the latest Prodigy version and that the string you've tested in spaCy directly also appears in your data?

Because the `textcat.teach` recipe prioritises the most uncertain predictions, it's possible that you won't see _all_ suggestions and _all_ pattern matches. But this shouldn't really be happening in the beginning. As a quick sanity check, you could also try running `ner.match` with your data and your patterns. This won't filter the incoming examples and just show you all pattern matches in your data.

---

<div class="post-metadata">

**Author:** ![bigbeaker](https://avatars.discourse-cdn.com/v4/letter/b/7feea3/32.png) [@bigbeaker](https://support.prodi.gy/u/bigbeaker)\
**Post date:** [January 7, 2019, 10:23am UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/6 "2019-01-07T10:23:13Z")

</div>

> [@ines](#):
>
> double-check you’re on the latest Prodigy version and that the string you’ve tested in spaCy directly also appears in your data?

I've tested the matching in Spacy and it works fine

I'm, on:  
Spacy version: 2.0.18 & Prodigy version: 1.6.1  
which are the latest

I also did get a couple of matches yesterday, but I had to go through ~500 tags to see two pattern matches.

> [@ines](#):
>
> As a quick sanity check, you could also try running `ner.match` with your data and your patterns.

Just tried this, working fine & picking up a pattern in every example.

So this must be something with how Prodigy is selecting the next example to show in textcat.teach as I'm expecting to see more pattern matches in the beginning of the tagging session which makes my session very unproductive as I have to reject a large amount of non-relevant examples.

Note: I'm using a custom language model & custom embeddings.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 7, 2019, 10:33am UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/7 "2019-01-07T10:33:51Z")

</div>

Thanks for the updates! In general, Prodigy will try to give you a good mix of pattern matches (especially in the beginning) and predictions, slowly focusing more on predictions than matches. But depending on the frequency of the matches and what the model is already predicting, it’s possible that this currently doesn’t always produce enough matches.

If you have a decent amount of patterns, maybe it makes sense to start off by annotating only matches and move on to annotating with a model in the loop later? You could use `ner.match` or build a slightly modified version that outputs tasks in the text classificatio style (with a label on top and no label next to the span). [See here](https://github.com/explosion/prodigy-recipes/blob/master/ner/ner_match.py) for a simplified version of the recipe, or check out `recipes/ner.py` in your Prodigy installation.

If you set `label_span=False` and `label_task=True` on the `PatternMatcher`, it’ll produce text classification tasks (top-level label, no label on the span):

```python
# Initialize the pattern matcher and load in the JSONL patterns
matcher = PatternMatcher(nlp, label_span=False, label_task=True).from_disk(patterns)

```

Make sure to also set `'view_id': 'classification'` to use the text classification interface.

---

<div class="post-metadata">

**Author:** ![bigbeaker](https://avatars.discourse-cdn.com/v4/letter/b/7feea3/32.png) [@bigbeaker](https://support.prodi.gy/u/bigbeaker)\
**Post date:** [January 7, 2019, 9:11pm UTC](https://support.prodi.gy/t/terms-teach-hangs-indefinitely-with-a-custom-word-vector-model/1106/8 "2019-01-07T21:11:19Z")

</div>

Oh great okay that seems to work well:

```python
@recipe('textcat.bootstrap',
        dataset=recipe_args['dataset'],
        spacy_model=recipe_args['spacy_model'],
        source=recipe_args['source'],
        api=recipe_args['api'],
        loader=recipe_args['loader'],
        patterns=recipe_args['patterns'],
        exclude=recipe_args['exclude'],
        resume=("Resume from existing dataset and update matcher accordingly",
                "flag", "R", bool))
def match(dataset, spacy_model, patterns, source=None, api=None, loader=None,
          exclude=None, resume=False):

    log("RECIPE: Starting recipe textcat.bootstrap", locals())
    DB = connect()
    # Create the model, using a pre-trained spaCy model.
    model = PatternMatcher(spacy.load(spacy_model), label_span=False, label_task=True).from_disk(patterns)
    log("RECIPE: Created PatternMatcher using model {}".format(spacy_model))
    if resume and dataset is not None and dataset in DB:
        existing = DB.get_dataset(dataset)
        log("RECIPE: Updating PatternMatcher with {} examples from dataset {}"
            .format(len(existing), dataset))
        model.update(existing)
    stream = get_stream(source, api=api, loader=loader, rehash=True,
                        dedup=True, input_key='text')
    return {
        'view_id': 'classification',
        'dataset': dataset,
        'stream': (eg for _, eg in model(stream)),
        'exclude': exclude
    }

```

But this doesn’t accept a label, so I’m assuming I have to run this first and then use textcat.teach to go over the examples collected here again
