# Disable sentence boundary detection in Spacy Parser

**URL:** <https://support.prodi.gy/t/disable-sentence-boundary-detection-in-spacy-parser/6083>\
**Category:** Uncategorized\
**Tags:** spacy\
**Created:** [November 6, 2022, 11:40pm UTC](https://support.prodi.gy/t/disable-sentence-boundary-detection-in-spacy-parser/6083 "2022-11-06T23:40:46Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![nlpfan](https://avatars.discourse-cdn.com/v4/letter/n/f0a364/32.png) [@nlpfan](https://support.prodi.gy/u/nlpfan)\
**Post date:** [November 6, 2022, 11:40pm UTC](https://support.prodi.gy/t/disable-sentence-boundary-detection-in-spacy-parser/6083/1 "2022-11-06T23:40:46Z")

</div>

My text input to Spacy is already in one sentence per line format. So I would like to switch off the sentence boundary detection in the parser.

Is there any config setting that controls the sentence boundary detection by the parser ? If not, is there a work around I can employ to let the parser assign dependency tags but not do sentence boundary detection ?

I would like to take advantage of the dependency tags generated by the parser so I believe excluding the parser from my pipeline is not the way to go.

Thanks!

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [November 7, 2022, 2:05pm UTC](https://support.prodi.gy/t/disable-sentence-boundary-detection-in-spacy-parser/6083/2 "2022-11-07T14:05:30Z")

</div>

I think there are two options here.

1. You could set up your own custom model, maybe using [pySBD](https://spacy.io/universe/project/python-sentence-boundary-disambiguation), and save that to disk. You can refer to this new saved model in your ner recipes.
2. You could write a custom recipe that takes care of the sentences in the loop. It might use something like:

```python
import srsly 

examples = srsly.read_jsonl("path/to/file.jsonl")

def sentence_stream(example):
    # Use your own split_sentence implementation here 
    for sentence in split_sentence(example['text']):
        yield {"text": sentence} 

stream = (sentence_stream(ex) for ex in examples)

```

Let me know if this doesn't work or if I'm misinterpreting your problem.

---

<div class="post-metadata">

**Author:** ![nlpfan](https://avatars.discourse-cdn.com/v4/letter/n/f0a364/32.png) [@nlpfan](https://support.prodi.gy/u/nlpfan)\
**Post date:** [February 19, 2023, 4:16pm UTC](https://support.prodi.gy/t/disable-sentence-boundary-detection-in-spacy-parser/6083/3 "2023-02-19T16:16:58Z")

</div>

I tried the first approach and it worked as expected. The text that I am working with is not well formed ( more like a bunch of sentence fragments, like text extracted from cells of a table) so dependency tags are not that useful.

So I am now using the balnk model ( blank:en) with ner.manual recipe and very satisfied with the results.

Thanks much for your help @koaning
