# Segmentation and newlines in ner.manual

**URL:** <https://support.prodi.gy/t/segmentation-and-newlines-in-ner-manual/494>\
**Category:** Uncategorized\
**Tags:** usage, ner, done\
**Created:** [April 17, 2018, 6:27am UTC](https://support.prodi.gy/t/segmentation-and-newlines-in-ner-manual/494 "2018-04-17T06:27:57Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [April 17, 2018, 5:25pm UTC](https://support.prodi.gy/t/segmentation-and-newlines-in-ner-manual/494/2 "2018-04-17T17:25:17Z")

</div>

Hi! I hope I’m understanding your question correctly – so you want to load in texts from a “custom” format and separate them into annotation tasks according to your own logic, right?

One option would of course be to pre-process your data, read in the input file, split on `\xa0` and then output a JSONL file with `{"text": "contract 1"}` etc per line.

You can also do this with a custom loader script in Python and then pip its output forward to the recipe. If no `source` argument is set on the command line, it will default to `stdin` (i.e. the output of the previous process). I’m describing this in more detail [on this thread](https://support.prodi.gy/t/does-prodigy-allow-loading-all-files-from-a-filepath/379/2?u=ines).

Here’s an example:

```python
import json

contracts = YOUR_LONG_TEXT.split('\xa0')
# you might also want to do some stripping of whitespace etc. here

for contract in contracts:
    task = {'text': contract}
    print(json.dumps(task)) # output dumped JSON

```

You can then pipe the tasks forward like this:

```bash
python your_script.py | prodigy ner.manual your_dataset en_core_web_sm --label SOME_LABEL

```

---

_[View the full topic](https://support.prodi.gy/t/segmentation-and-newlines-in-ner-manual/494)._
