# Wrap breaks for long documents

**URL:** <https://support.prodi.gy/t/wrap-breaks-for-long-documents/6891>\
**Category:** Uncategorized\
**Tags:** ner, custom, relations\
**Created:** [November 10, 2023, 2:37pm UTC](https://support.prodi.gy/t/wrap-breaks-for-long-documents/6891 "2023-11-10T14:37:02Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Martin](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/martin/32/4214_2.png) [@Martin](https://support.prodi.gy/u/Martin)\
**Post date:** [November 10, 2023, 2:37pm UTC](https://support.prodi.gy/t/wrap-breaks-for-long-documents/6891/1 "2023-11-10T14:37:02Z")

</div>

Hi everyone,

I have a specific need to annotate quite long documents with relations and entities, also showing some metadata, and we have begun work on a custom recipe to do that. We want to have the full document in our browser (scrolling would not be an issue) but the wrap for line breaks in the `relation` view seems to break past 7'000 characters (I don't have the token count yet, sorry), giving us :

 ![prodigy-error](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/9/92cf4f11e35a52a8312cbee4acadf3abd9c47015.png)

For reference, here is our page :

 ![prodigy-working](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/f/fa40be965770ef9529625e1bee56b3d307bcb7e9.png)

Here is the recipe definition, nothing spectacular :

```python
{
        "dataset": dataset,
        "view_id": "blocks",
        "stream": stream,
        "config": {
            "lang": lang,
            "relations_span_labels": [
                "Avocat",
                "Cabinet",
                "Partie Personne Morale",
                "Forme Juridique",
            ],
            "labels": ["Substitue", "Représente", "A pour status", "membre de"],
            "blocks": [
                {
                    "view_id": "html",
                    "html_template": "<div> Annoter la decision suivante: {{stream['meta']['short_title']}} </div>",
                },
                {"view_id": "relations"},
            ],
        },
    }

```

I would love at least a work around to have line breaks in our interface, as we have line breaks in our documents, and they are significant for our purposes. I have seen [newlines in relations annotation](https://support.prodi.gy/t/newlines-in-relations-annotation/3524) and it doesn't specifically help given that wrap breaks on us ! Is there any way to break lines on newline tokens ?

As an aside, I noticed that on wider screens, prodigy uses very little of the available space, which is a bummer for us given the quantity of text needed to grasp our documents (we work on legal decisions in France).

Edit: For readability, here is the trace given by the error :

```python
[Exception... "Failure" nsresult: "0x80004005 (NS_ERROR_FAILURE)" location: "JS frame :: http://localhost:8080/bundle.js :: setHeight :: line 170" data: no]
StageWrap@http://localhost:8080/bundle.js:220:11863
FiberProvider@http://localhost:8080/bundle.js:220:10007
div
div
_Content@http://localhost:8080/bundle.js:164:7859
bundle.js/createWithStyles/</WithStyles<@http://localhost:8080/bundle.js:88:15449
_Relations@http://localhost:8080/bundle.js:220:58610
bundle.js/createWithStyles/</WithStyles<@http://localhost:8080/bundle.js:88:15449
_Blocks@http://localhost:8080/bundle.js:164:20801
bundle.js/createWithStyles/</WithStyles<@http://localhost:8080/bundle.js:88:15449
div
_CardWithoutTheme@http://localhost:8080/bundle.js:220:89391
bundle.js/createWithStyles/</WithStyles<@http://localhost:8080/bundle.js:88:15449
Card@http://localhost:8080/bundle.js:220:92095
ErrorBoundary@http://localhost:8080/bundle.js:220:95619
div
div
_Annotator@http://localhost:8080/bundle.js:220:100285
bundle.js/createWithStyles/</WithStyles<@http://localhost:8080/bundle.js:88:15449
ConnectFunction@http://localhost:8080/bundle.js:71:12533
main
div
_class2@http://localhost:8080/bundle.js:93:42083
_Main@http://localhost:8080/bundle.js:220:118586
bundle.js/createWithStyles/</WithStyles<@http://localhost:8080/bundle.js:88:15449
ConnectFunction@http://localhost:8080/bundle.js:71:12533
ThemeProvider3@http://localhost:8080/bundle.js:75:24970
App@http://localhost:8080/bundle.js:222:5900
Provider@http://localhost:8080/bundle.js:75:1139

```

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [November 10, 2023, 3:47pm UTC](https://support.prodi.gy/t/wrap-breaks-for-long-documents/6891/2 "2023-11-10T15:47:26Z")

</div>

hi @Martin,

Thanks for your post.

Unfortunately, it seems like you're getting this error due to the long size of the document, especially for relations:

> [@relation is responding very slowly](https://support.prodi.gy/t/relation-is-responding-very-slowly/4732/2):
>
> Hi and sorry about this – it's currently expected that the interface becomes less performant for very long documents with many tokens, and we're working on a rewrite that doesn't have this problem. I think what makes it additionally tricky in your case is that you just end up with more tokens overall, due to the way the characters are segmented. As a workaround, one thing you can do is use patterns to disable all tokens that you know won't ever be part of a relation – of course, only if that's …

One option could be to disable tokens (use `--disable-patterns`) you're not interested in, for example see the [docs](https://prodi.gy/docs/dependencies-relations#custom).

However, I suspect your document is still way too long.

> I would love at least a work around to have line breaks in our interface, as we have line breaks in our documents, and they are significant for our purposes. I have seen [newlines in relations annotation](https://support.prodi.gy/t/newlines-in-relations-annotation/3524) and it doesn't specifically help given that wrap breaks on us ! Is there any way to break lines on newline tokens ?

Why don't you write a pre-processing script for this?

For example, this [Stack Overflow](https://stackoverflow.com/a/75237658) shows how to break up by new line characters this before using spaCy. You could combine this with this example:

> [@how to annotate a longer text in rel.manual?](https://support.prodi.gy/t/how-to-annotate-a-longer-text-in-rel-manual/6095/2):
>
> Is there a reason why you can't split the data into smaller segments? I can imagine that long pieces of text require a lot of scrolling and can make it much harder to annotate. I sometimes like to use spaCy to preprocess my data, via something like: import spacy nlp = spacy.load("en\_core\_web\_md") # I'm assuming a list/generator of texts loaded in Python orig\_examples = ... # This function can then turn them into seperate json blobs with sentences def to\_sentences(orig\_examples): for d…

So now your source (input) file would be text broken up by the new line character, so you can still use the `rel.manual` recipe.

> [@Martin](#):
>
> As an aside, I noticed that on wider screens, prodigy uses very little of the available space, which is a bummer for us given the quantity of text needed to grasp our documents (we work on legal decisions in France).

You can modify the card size by [modifying the CSS](https://prodi.gy/docs/custom-interfaces#js-css) in your [configuration](https://prodi.gy/docs/install#config) like:

```python
"global_css": ".prodigy-container { max-width: 950px; }"

```

There are also other options (e.g., change font size) in the [docs](https://prodi.gy/docs/named-entity-recognition#long-text) to modify the UI space.

Hope this helps!

---

<div class="post-metadata">

**Author:** ![Martin](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/martin/32/4214_2.png) [@Martin](https://support.prodi.gy/u/Martin)\
**Post date:** [November 14, 2023, 9:58am UTC](https://support.prodi.gy/t/wrap-breaks-for-long-documents/6891/3 "2023-11-14T09:58:28Z")

</div>

So, after a bit of experimentations, we settled (for now) on a sliding window of around 10 lines of text, to annotate our documents. Line breaks in our documents are significant so we cannot just throw them away, and there is an issue on relations. Some relations are between entities that are really far apart (sometimes up to 5 lines) which means we cannot annotate all relations in one sliding window as some might be outside the example while others are present and must be annotated...

I don't know whether to mark examples where we only have some values and relations but not all of them as rejected examples, or if we should annotate them anyway, even if they are missing relations.

I foresee some issues trying to train a relation model with all these separate annotation of 10 lines windows, given that on some windows we will have relations with out of the window entities... I am not sure about how to make this workflow better. Do you have any idea ? I'm thinking about maybe trying to add a pseudo-sentencizer that separates blocks of significant value (where the relations are sure to be) but it's a lot of extra work...

---

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [November 16, 2023, 9:57am UTC](https://support.prodi.gy/t/wrap-breaks-for-long-documents/6891/4 "2023-11-16T09:57:37Z")

</div>

Hi @Martin,

My colleague @ryanwesslen is OOO for a few days so I thought I would jump in to provide some advice.  
I do share your concern that "incomplete" examples will affect negatively the performance of the classifier. It is also true, though, that loading full documents is not a good annotator experience either even if it was technically possible (and bad UI usually translates into inaccurate labels).

Just to clarify: by sliding window you mean that you have window of size 10 (lines) and you move it by increments of \<=5, which effectively means that you look at each line multiple times in different contexts, right?  
In that case, given that most relations are within 5 lines you should be able to see them all at some point.  
Assuming that this is the setup (as opposed to splitting the document into chunks size 10 and iterating through that), it's probably best to annotate only complete examples and have a postprocessing step to merge all the relations. This way you should be able to catch them all or at least catch high enough number and variation so that the model will be able to generalize.  
In that case, for the efficiency of the annotation process we would recommend performing NER and REL separately. NER without the the sliding window and REL with the sliding window over the NER dataset as discussed above.
