# Segmenting examples with long spans as NERs

**URL:** <https://support.prodi.gy/t/segmenting-examples-with-long-spans-as-ners/652>\
**Category:** Uncategorized\
**Tags:** ner\
**Created:** [June 27, 2018, 1:28pm UTC](https://support.prodi.gy/t/segmenting-examples-with-long-spans-as-ners/652 "2018-06-27T13:28:50Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![honnibal](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/honnibal/32/35_2.png) [@honnibal](https://support.prodi.gy/u/honnibal)\
**Post date:** [June 27, 2018, 11:06pm UTC](https://support.prodi.gy/t/segmenting-examples-with-long-spans-as-ners/652/2 "2018-06-27T23:06:44Z")

</div>

Hi @pvcastro,

As I mentioned in the last thread, I'm suspicious of using the entity recognised for these long spans. I think you should try applying sentence labels, and perhaps also marking words which are important for the category you're interested in. Then you can use the dependency parse to find the claim boundaries. You can find documentation about the dependency parser here: [Linguistic Features · spaCy Usage Documentation](https://spacy.io/usage/linguistic-features#section-dependency-parse)

> [@pvcastro](#):
>
> But my actual problem is that if I don’t use the ‘unsegmented’ parameter, the recipe is losing the spans for each example after segmenting them. My dataset of 400 examples goes to a few thousand sentences, but only about 11 of these have the original spans, all the other thousands have empty spans. Do you know what could be the problem?

There should be as many spans, whether you set unsegmented or not --- unless the spans cross segmentation boundaries. This sounds like it might be a bug; we'll look into it. At first glance the segmentation function looks correct, and it's passing our tests. But I'll play around with your sample and see if I can find the problem.

---

_[View the full topic](https://support.prodi.gy/t/segmenting-examples-with-long-spans-as-ners/652)._
