# Annotating dependecies for very long sentences

**URL:** <https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024>\
**Category:** Uncategorized\
**Tags:** usage, relations\
**Created:** [March 14, 2021, 6:07pm UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024 "2021-03-14T18:07:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![APagano](https://avatars.discourse-cdn.com/v4/letter/a/3bc359/32.png) [@APagano](https://support.prodi.gy/u/APagano)\
**Post date:** [March 14, 2021, 6:07pm UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/1 "2021-03-14T18:07:41Z")

</div>

We have not yet subcribed to Prodigy and wonder whether Prodigy will be useful for our annotation project of dependency relations. Our corpus has very long sentences and no web tool or interface has yet proved helpful for that. We have tried Arborator and several others, but dragging arcs is impossible when the sentence is very long. Working on the CoNNLU file is the only way out so far, but it sometimes makes you annotate the wrong numbers. Any insight as to whether Prodigy might be the solution? Thanks for feedback on this.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 16, 2021, 3:17am UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/2 "2021-03-16T03:17:21Z")

</div>

Hi! It's definitely true that once you have long texts with lots of very long dependencies, the visual gain you get from drawing arcs on top of the text can be low and allowing every token to be connected to every other token can make things _more_ complex instead of solving complexity with a visual UI. This is more of a conceptual problem and the best solution ultimately depends on the type of annotation and what the dependencies represent.

When you say long sentences, how long are they on average? And what type of dependencies are you annotating? Are you working with syntax where you essentially need to connect every token to a root, or are you labelling different and more sparse annotations?

If you're _not_ annotating syntax, there are various things you can do to reduce the complexity – for example, [disabling tokens you don't need](https://prodi.gy/docs/dependencies-relations#coref) automatically (e.g. for coref) or annotating [abstract representations](https://support.prodi.gy/t/using-relations-interface-for-large-texts/3461/4) (e.g. for sentence alignment).

If you _are_ annotating syntactic dependencies (and especially if your goal is to create a proper treebank), there's obviously no way around labelling every token. I still think Prodigy can be useful here: you'll be able to toggle between line wrapping and inline view, hide/show the arcs to get a better overview and you can assign dependencies by clicking (instead of dragging). You can see a minimal example of a short sentence [here](https://prodi.gy/demo?view_id=dep), but the experience will be the same for long sentences. (Where it could become a bit trickier to ensure high performance is if your sentences are longer than ~300 tokens on average – but that would be _very_ long sentences, so I assume yours are a bit shorter than that?)

Btw, I saw you're also emailed, and we're happy to set you up with an adademic license so you can try it out 🙂

---

<div class="post-metadata">

**Author:** ![APagano](https://avatars.discourse-cdn.com/v4/letter/a/3bc359/32.png) [@APagano](https://support.prodi.gy/u/APagano)\
**Post date:** [March 16, 2021, 1:12pm UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/3 "2021-03-16T13:12:21Z")

</div>

Many thanks indeed for your reply! I am annotating syntax dependencies in clinical text. Long sentences have on average 150 tokens, with plenty of list and conj relations, which demands drawing long arcs.  
How about the output file once we annotate syntax dependencies? Can we export data as txt (CoNLL)? json? csv?  
I have already applied for an academic license. I hope I am able to install it at all (I am a linguist, not too computer savvy).  
Anyway, I look forward to trying Prodigy. Regards.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 17, 2021, 12:22am UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/4 "2021-03-17T00:22:41Z")

</div>

> [@APagano](#):
>
> How about the output file once we annotate syntax dependencies? Can we export data as txt (CoNLL)? json? csv?

Prodigy's default output format is JSON and you can find an example of the dependencies/relations output here: [https://prodi.gy/docs/api-interfaces#relations](https://prodi.gy/docs/api-interfaces#relations) As you can see from the data, it should have everything you need: each dependency with its head, child and label, as well as the reference to the head and child tokens. If you need a different format, you should be able to convert it.

> [@APagano](#):
>
> I have already applied for an academic license. I hope I am able to install it at all (I am a linguist, not too computer savvy).

Sure, I totally understand! There are a few things about setting up a Python development environment that can be a bit tricky or unintuitive if you haven't done this before. But in general, we try to work with standard technologies and concepts wherever possible, so if you have a colleague who has experience with Python etc., they should be able to help without having to learn anything super specific to Prodigy 🙂

---

<div class="post-metadata">

**Author:** ![APagano](https://avatars.discourse-cdn.com/v4/letter/a/3bc359/32.png) [@APagano](https://support.prodi.gy/u/APagano)\
**Post date:** [March 17, 2021, 12:12pm UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/5 "2021-03-17T12:12:22Z")

</div>

After hours of trial and error, I managed to install Prodigy on my macbook. But now I am lost. I got "No module named prodigy"

Adrianas-MacBook-Pro:~ ariadna$ python3 -m prodigy  
Traceback (most recent call last):  
File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/runpy.py", line 188, in \_run\_module\_as\_main  
mod\_name, mod\_spec, code = \_get\_module\_details(mod\_name, \_Error)  
File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/runpy.py", line 147, in \_get\_module\_details  
return \_get\_module\_details(pkg\_main\_name, error)  
File "/Library/Frameworks/Python.framework/Versions/3.9/lib/python3.9/runpy.py", line 111, in \_get\_module\_details  
**import** (pkg\_name)  
File "/Users/ariadna/Library/Python/3.9/lib/python/site-packages/prodigy/ **init**.py", line 1, in   
from .util import init\_package  
ModuleNotFoundError: No module named 'prodigy.util'

Besides that, how do I get to an interface? Tks for any help.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 18, 2021, 1:00am UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/7 "2021-03-18T01:00:34Z")

</div>

Hi! Sorry you were having trouble! The latest stable Prodigy installer currently supports Python 3.6, 3.7 and 3.8 but it looks like you're running 3.9. So this is likely the problem here. So you can either install Python 3.8, or use a tool like [`pyenv`](https://github.com/pyenv/pyenv) that lets you run multiple versions of Python.

We'll be adding wheels for 3.9 to the upcoming [new version](https://support.prodi.gy/t/prodigy-nightly-spacy-v3-support-ui-for-overlapping-spans-improved-feeds-more/3861/) – we first had to wait for all our dependencies to support it before we could build a version of Prodigy for 3.9.

> [@APagano](#):
>
> Besides that, how do I get to an interface? Tks for any help.

Once you're set up, you can start an annotation workflow, also called "recipe" from the command line. Different recipes have different commands and options.

The getting started guide might be a good place to start and it explains the most important concepts: [Prodigy 101 – everything you need to know · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs#first-steps)

---

<div class="post-metadata">

**Author:** ![APagano](https://avatars.discourse-cdn.com/v4/letter/a/3bc359/32.png) [@APagano](https://support.prodi.gy/u/APagano)\
**Post date:** [March 18, 2021, 4:05pm UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/8 "2021-03-18T16:05:19Z")

</div>

Thanks again! I tried one of the recipes but got  
Can't read file: disable\_patterns.jsonl  
Is there a step by step document/tutorial?

Sorry about this.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 19, 2021, 1:02am UTC](https://support.prodi.gy/t/annotating-dependecies-for-very-long-sentences/4024/9 "2021-03-19T01:02:18Z")

</div>

> [@APagano](#):
>
> Thanks again! I tried one of the recipes but got  
> Can't read file: disable\_patterns.jsonl  
> Is there a step by step document/tutorial?

The `--disable-patterns` argument lets you define a path to a patterns file to disable tokens you know are unselectable. It's an optional setting, so you don't have to use it if it's not relevant for your use case.

The following command will start the [`rel.manual`](https://prodi.gy/docs/recipes#relations) workflow, save the annotations to the dataset `your_datset`, pre-tokenize the text with English tokenization rules, load in a `.txt` file with raw text (replace this with your texts) and let you assign the labels `A`, `B` and `C`:

```bash
prodigy rel.manual your_dataset blank:en ./path/to/data.txt --label A,B,C

```

You can also load in your text as JSON(L) or CSV instead. You can also load in pre-tokenized data by providing a list of `"tokens"` ([see here](https://prodi.gy/docs/api-interfaces#relations) for the format).
