# How to extract dependencies in spaCy after using prodigy rel.manual?

**URL:** <https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134>\
**Category:** Uncategorized\
**Tags:** usage, spacy, relations\
**Created:** [April 13, 2021, 7:51pm UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134 "2021-04-13T19:51:34Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![jamiehannaford](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/jamiehannaford/32/2093_2.png) [@jamiehannaford](https://support.prodi.gy/u/jamiehannaford)\
**Post date:** [April 13, 2021, 7:51pm UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/1 "2021-04-13T19:51:34Z")

</div>

👋 I'm trying to develop a model using NER and relation extraction with prodigy and had a usage question.

To start off, I generated a JSONL dataset that contained a few thousand sentences and were pre-labelled with the 3 spans I want to relate together (I retrieved their offsets using `PhraseMatcher`). I then ran:

```python
$ prodigy rel.manual ner_exp_restr_dep en_core_web_lg ./output.jsonl \
   --label HAS_COSTS,IN_YEAR \
   --span-label EXPENSE,MONEY,DATE \
   --add-ents \
   --wrap

```

and spent about 30 minutes annotating 100 or examples with relation data. The basic idea is that a `EXPENSE` span relates to a `MONEY` span, which relates to a `DATE` span. After saving this to the DB, I ran:

```python
$ prodigy train rel en ner_exp_restr_dep

```

which exported to a local directory. I then imported the model with:

```python
import spacy

nlp = spacy.load("./rel/model-last")

doc = nlp("In 2020 we recorded $20 million in impairment charges")

for ent in doc.ents:
    print(ent.text, ent.label_)

# 2020 DATE
# $20 million MONEY
# impairment charges EXPENSE

```

there doesn't seem to be a way to map from `EXPENSE` -\> `MONEY` -\> `DATE`.

How do I map from one entity to another using the relations extracted in prodigy? I didn't see anything in the docs about the next steps required.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [April 14, 2021, 1:37am UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/2 "2021-04-14T01:37:24Z")

</div>

Hi! I'm kinda surprised `train rel` worked and didn't raise an error because there's no component to train here 🤔 Actually, I just remembered that you were the same person who disabled `use_plac` as a workaround on [this thread](https://support.prodi.gy/t/argumenterror-for-prodigy-train-on-v1-11-0a6/4132/2) – this is likely the problem because it skips CLI argument validation. So I'd recommend turning that back on and using the other workaround I provided in the thread.

There's currently no built-in component for relation extraction in spaCy, so you will have to use your own implementation, depending on how you want your relation extraction to work. Here are some related threads:

> [@Relationship between named entities](https://support.prodi.gy/t/relationship-between-named-entities/3171):
>
> I'm having a lot of fun exploring prodigy for a text mining task. Training a NER model using annotations is extremely rewarding, and works surprisingly well with only a small amount of manual annotations. So far, I've been using prodigy train ner ... for assigning entities. Let's say I have two entity types, FIRM and TECH. Some data to illustrate my task: [IBM][FIRM] has identified [hybrid cloud][TECH] as the growth area it will focus on. In early March 2017, [Snapchat][FIRM] officially went…

> [@How to train and correct a Named Entity Recognition with relation extraction](https://support.prodi.gy/t/how-to-train-and-correct-a-named-entity-recognition-with-relation-extraction/3724):
>
> I can not find how to train an entity and relation annotation done at the same time. And how to correct one time done? I mean if I do prodigy rel.manual ner\_rels blank:en ./data.jsonl --label SUBJECT,LOCATION --span-label PERSON,GPE --wrap --add-ents and then I use prodigy train parser ner\_rels en\_core\_web\_lg --output ./model The nerd and relation model are trained at the same time? ANd how I correct the results like ner.correct but with relations too? thx

> [@Building model from annotated directional relations and dependencies dataset](https://support.prodi.gy/t/building-model-from-annotated-directional-relations-and-dependencies-dataset/3700):
>
> hello, I used this command to annotate directional relations and dependencies between tokens: prodigy rel.manual my\_dataset my\_pre\_annotated\_model ./file\_to\_data.jsonl --label label1,label1 --span-label ner\_label1,ner\_label2,ner\_labe3 --wrap after that, how can i convert the resulted dataset to a model so i can correct it and increase its accuracy? i tried this: prodigy dep.batch-train my\_dataset blank:en --output ./model\_Dependency --label label1,label2 --eval-split 0.2 --n-iter 10 but i …

For spaCy v3, @SofieVL recorded this in-depth tutorial on how to implement an entity relation extraction component from scratch. The code for this is available as a spaCy project so you can experiment with it. Even if you want to do something more custom, the video has a lot of helpful pointers on how to model the problem:

[![](https://img.youtube.com/vi/8HL-Ap5_Axo/hqdefault.jpg "SPACY v3: Custom trainable relation extraction component") ](https://www.youtube.com/watch?v=8HL-Ap5_Axo)

---

<div class="post-metadata">

**Author:** ![jamiehannaford](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/jamiehannaford/32/2093_2.png) [@jamiehannaford](https://support.prodi.gy/u/jamiehannaford)\
**Post date:** [April 14, 2021, 8:57pm UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/3 "2021-04-14T20:57:03Z")

</div>

Thanks @ines, those links were super helpful. I tried cloning that `tutorials/rel_component` project, and then ran:

```python
spacy project run all

```

to process the data and train an initial model. But when I ran this script from the same directory (using the first sentence in `assets/annotations.jsonl`), it didn't find any of `doc._.rel` fields:

```python
import spacy
from scripts.rel_pipe import *
from scripts.rel_model import *

nlp = spacy.load("training/model-best")

doc = nlp("Furthermore, Smad-phosphorylation was followed by upregulation of Id1 mRNA and Id1 protein, whereas Id2 and Id3 expression was not affected.")

print("spans", [(e.start, e.text, e.label_) for e in doc.ents])

for value, rel_dict in doc._.rel.items():
    print(f"{value}: {rel_dict}")

# ℹ Could not determine any instances in doc - returning doc as is.
# spans []

```

I tried instantiating a `Doc` object directly (like in `evaluate.py`), but it doesn't help either:

```python
words = ['Luciferase', 'assays', 'revealed', 'a', 'approximately20-fold', 'increased', 'transcriptional', 'activity', 'of', 'the', '1025', 'bp', 'sequence', 'as', 'compared', 'to', 'the', 'empty', 'vector', ',', 'indicating', 'that', 'we', 'had', 'identified', 'an', 'active', 'A3', 'G', 'promoter', 'sequence', '(', 'Figure', '3B', ')', '.'] 
spaces = [' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', '', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', '', ' ', ' ', ' ', '', ' ', '', '', '']
doc = Doc(nlp.vocab, words, spaces)

# spans []
# 

```

I'm wondering if I'm missing something basic here 🤔 The model seems trained, I'm importing it with `spacy.load` but it's not finding any of the labels or entities.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [April 15, 2021, 8:28am UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/4 "2021-04-15T08:28:37Z")

</div>

> [@jamiehannaford](#):
>
> I'm wondering if I'm missing something basic here 🤔 The model seems trained, I'm importing it with `spacy.load` but it's not finding any of the labels or entities.

Does your model have an entity recognizer? The relation extraction component requires named entities and in the tutorial, Sofie uses gold-standard entities as the input for simplicity. But if your `doc` doesn't have any `doc.ents`, the relation extraction won't have any entities to choose from and predict over.

---

<div class="post-metadata">

**Author:** ![jamiehannaford](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/jamiehannaford/32/2093_2.png) [@jamiehannaford](https://support.prodi.gy/u/jamiehannaford)\
**Post date:** [April 15, 2021, 4:09pm UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/5 "2021-04-15T16:09:00Z")

</div>

Ah, I see. I had assumed that both the entities and the relations were defined in `assets/annotations.jsonl` and that running `project run all` would handle both the entity _and_ relation extraction learning. Does it only do the former?

If so, is there a way to plug in my own entity recognizer model here? I'm hoping that if the relation extractor is generic enough, I could just bring my own NER model and JSONL file, and stuff will appear in `doc._.rels`.

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [April 15, 2021, 11:36pm UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/6 "2021-04-15T23:36:58Z")

</div>

Hi Jamie,

You could in principle train a blank NER model and REL model from scratch within the same pipeline, by adding an ner component to your training config.

A word of warning though. If your annotation has been focusing on getting the relations right, this might not be the best data to train the NER model on. Ideally, you'd want the NER model to be generic enough to pick up all mentions of the entities you're interested in - not just those that are also expressed as being in a relation.

Because you annotated the dataset with `rel.manual`, you'll only be presented with sentences that have at least 2 entities in them, because a relation isn't possible otherwise. This might bias your NER model if you're only training on the entities in these sentences. This is why it might often make sense to train your NER separately from your REL model.

---

<div class="post-metadata">

**Author:** ![jamiehannaford](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/jamiehannaford/32/2093_2.png) [@jamiehannaford](https://support.prodi.gy/u/jamiehannaford)\
**Post date:** [April 16, 2021, 5:48pm UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/7 "2021-04-16T17:48:07Z")

</div>

@SofieVL That's great context, thank you! I'll give it a go.

One last question: is it possible for the relation extraction process to understand and parse relations for different types of entity? Or is it best for the model to be trained against a single named entity? I'll give an example:

> I paid $2 (MONEY) for an apple (FRUIT) yesterday (DATE)

> I paid $100 (MONEY) for Lego (TOY) last week (DATE)

Although `FRUIT` and `TOY` are different entities (and will have their own NER models), their relationship to other gold-standard entities (in this case `MONEY` and `DATE`) are the same.

I'm thinking I could use prodigy to generate a JSONL file that references both types of entities and use `spacy project run all` against that, but didn't know if adding more label types would reduce the accuracy.

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [April 19, 2021, 6:48am UTC](https://support.prodi.gy/t/how-to-extract-dependencies-in-spacy-after-using-prodigy-rel-manual/4134/8 "2021-04-19T06:48:03Z")

</div>

It depends on the exact implementation of your relation extraction component, but in general I think it would be beneficial to learn both these examples at the same time, even if they pertain to different entity types.

The experimental REL code that Ines linked earlier in this thread, doesn't currently take the entity types as features for relation extraction, so that shouldn't be a problem. The semantics of the relationship is the same, which is the most important bit for Machine Learning.
