# Tip: Preprocessing text (whitespace, unicode) with textacy

**URL:** <https://support.prodi.gy/t/tip-preprocessing-text-whitespace-unicode-with-textacy/343>\
**Category:** Uncategorized\
**Tags:** usage, custom, solved\
**Created:** [February 26, 2018, 2:57pm UTC](https://support.prodi.gy/t/tip-preprocessing-text-whitespace-unicode-with-textacy/343 "2018-02-26T14:57:31Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 26, 2018, 2:57pm UTC](https://support.prodi.gy/t/tip-preprocessing-text-whitespace-unicode-with-textacy/343/1 "2018-02-26T14:57:31Z")

</div>

Inspired by [this discussion](https://github.com/chartbeat-labs/textacy/issues/168), I wrote a little Prodigy recipe that preprocesses a stream of text using the [`textacy` package](https://github.com/chartbeat-labs/textacy) for higher-level NLP with spaCy.

The recipe takes an input source and can remove double or trailing whitespace, fix broken unicode and mojibake, convert non-ASCII characters to the closest ASCII characters and replace accented characters with unaccented. There are [various other parameters](https://chartbeat-labs.github.io/textacy/api_reference.html#textacy.preprocess.preprocess_text) available as well, which you can plug in in a similar way.

⚠️ Disclaimer: If you want the model to learn how to deal with unclean text, it also needs to see examples of this during training. So in cases like this, it’s usually not recommended to clean up your training data (actually, quite the opposite – but that’s something for another data augmentation recipe).

To use the recipe, you need to install `textacy`:

```bash
pip install textacy

```

Then place the following in a recipe file, e.g. `recipe.py`:

```python
import prodigy
from textacy.preprocess import normalize_whitespace, preprocess_text
import ujson

@prodigy.recipe('preprocess',
    source=prodigy.recipe_args['source'],
    normalize_ws=('Normalize whitespace', 'flag', 'ws', bool),
    fix_unicode=('Fix broken unicode', 'flag', 'u', bool),
    transliterate=('Convert non-ASCII if possible', 'flag', 't', bool),
    no_accents=('Replace accented characters with unaccented', 'flag', 'na', bool))
def preprocess(source, normalize_ws=False, fix_unicode=False, 
               transliterate=False, no_accents=False):
    stream = prodigy.get_stream(source)
    for eg in stream:
        text = eg['text']
        if normalize_ws:
            text = normalize_whitespace(text)
        text = preprocess_text(text, fix_unicode=fix_unicode,
                               transliterate=transliterate, no_accents=no_accents)
        eg['text'] = text
        # write example to stdout
        print(ujson.dumps(eg, escape_forward_slashes=False, ensure_ascii=False))

```

For an overview of the available command-line options, you can run:

```bash
prodigy preprocess --help -F recipe.py

```

You can preview the preprocessed stream like this:

```bash
prodigy preprocess your_data.jsonl -ws -u -na -F recipe.py | less

```

And then pipe it forward to another recipe – for example:

```bash
prodigy preprocess your_data.jsonl -ws -u -na -F recipe.py | ner.teach your_dataset en_core_web_sm

```

---

<div class="post-metadata">

**Author:** ![amoux](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/amoux/32/734_2.png) [@amoux](https://support.prodi.gy/u/amoux)\
**Post date:** [November 4, 2019, 8:58am UTC](https://support.prodi.gy/t/tip-preprocessing-text-whitespace-unicode-with-textacy/343/2 "2019-11-04T08:58:01Z")

</div>

Hello! I am currently working with social-media text, and I was questioning if the following preprocessing is the "correct-practice." I want to increase the model's accuracy in identifying various classes from the textcat annotations. So I have avoided doing any preprocessing that could modify and change the original form of the text. Here is a sample of how most of all the documents look like before and after minor processing. Should I get rid of the newlines and extra spacing?

- Before preprocessing:

```python
raw_docs = [
'next video : bill guesses how much luxurious cars costs',
'Good night now,I am really tired so I am going to bed\nGood luck with your next video 💤💤💤💤😴😴😴😴5mins'
]

```

- After preprocessing ( spacy's sentence tokenizer and encoding ASCII ):

```python
sents = [
'next video : bill guesses how much luxurious cars costs',
'Good night now,I am really tired',
'so I am going to bed\n',
'Good luck with your next video 5mins'
]

```

* * *

- update

> I watched the FAQ video where the topic _ **"What if I need to label long texts"** _ is discussed. I found the information I needed! This tool has made experimenting a lot easier, thank you to all at spacy/prodigy 🙂

---

<div class="post-metadata">

**Author:** ![honnibal](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/honnibal/32/35_2.png) [@honnibal](https://support.prodi.gy/u/honnibal)\
**Post date:** [November 7, 2019, 12:02pm UTC](https://support.prodi.gy/t/tip-preprocessing-text-whitespace-unicode-with-textacy/343/3 "2019-11-07T12:02:45Z")

</div>

Glad to hear you solved the problem!

One thing you might try as well is a little extra normalization. Specifically, normalizing the punctuation a bit might help slightly as well, depending on the specifics of your data. It's less important for text classification, but if you're working with the other models, it might help.
