# Arabic text rendering in ner.manual

**URL:** <https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253>\
**Category:** Uncategorized\
**Tags:** front-end, solved\
**Created:** [January 27, 2018, 5:37pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253 "2018-01-27T17:37:41Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![andy](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/andy/32/966_2.png) [@andy](https://support.prodi.gy/u/andy)\
**Post date:** [January 27, 2018, 5:37pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253/1 "2018-01-27T17:37:41Z")

</div>

I’m running into an issue where Arabic text is not being correctly rendered in `ner.manual`. Words are spread across a couple lines above and below, and then jump down into their correct place once highlighted (see example below). The word that was above the line is now incorporated into it.

 ![22 PM](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/1X/01b635037dfa9c3d45da83207c2c08999bf47555.jpg)

 ![59 PM](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/1X/125680d27ba8fe095ebc70282eef1fec9bdbd290.jpg)

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 27, 2018, 5:54pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253/2 "2018-01-27T17:54:36Z")

</div>

Thanks a lot for the report!

Do you have one or two example texts that I can copy-paste to debug the interface? And could you check the `"tokens"` property on the incoming tasks that are created and verify that it’s all split correctly? Just so we can make sure that spaCy’s tokenization is not the problem here.

~~(I was going to suggest modifying the writing direction to RTL, but I just realised that the manual UI doesn’t support the `"card_css"` config setting at the moment, since this can easily lead to unexpected results in this particular interface. But if you want to hack at it in your browser’s dev tools, I’d be curious to see the effect of setting `direction: rtl` on the parent container’s styles.)~~

Sorry, I was wrong: `"card_css": "direction: rtl"` is definitely possible, even in manual NER mode. Not sure if this is the full solution, though.

---

<div class="post-metadata">

**Author:** ![andy](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/andy/32/966_2.png) [@andy](https://support.prodi.gy/u/andy)\
**Post date:** [January 29, 2018, 9:48pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253/3 "2018-01-29T21:48:57Z")

</div>

Adding `"card_css": "direction: rtl"` to `.prodigy.json` fixed the weird display problem. Thank you!

There’s a separate tokenization problem, which is that spaCy is including periods in with the previous token in the `xx` and `en` models. But that’s an Arabic in spaCy issue, not a Prodigy issue.

If you still want a piece of example text, here’s one:

ﺕﻮﺟﺩ ﺏﺎﻠﺒﺣﺮﻴﻧ ﺄﻜﺑﺭ ﻢﻘﺑﺭﺓ ﺕﺍﺮﻴﺨﻳﺓ ﺏﺎﻠﻋﺎﻠﻣ ﻮﻬﻳ ﻊﻟﻯ ﺶﻜﻟ ﺭﻭﺎﺒﻳ (ﺕﻼﻟ) ﻮﺘﺴﻣﻯ ﻢﻗﺎﺑﺭ ﻉﺎﻠﻳ ﻢﺨﺘﻠﻓﺓ ﺎﻠﺤﺠﻣ ﻭﺎﻠﺸﻜﻟ. ﻮﺗﻮﺟﺩ ﺐﻤﻨﻄﻗﺓ ﻉﺎﻠﻳ ﺏﺎﻠﻤﺣﺎﻔﻇﺓ ﺎﻟﻮﺴﻃﻯ، ﻮﻠﻫﺬﻫ ﺎﻠﻤﻗﺎﺑﺭ ﺕﺍﺮﻴﺧ ﻉﺮﻴﻗ ﻕﺪﻴﻣ ﻢﻧﺫ ﻊﻫﺩ ﺢﺿﺍﺭﺓ ﺪﻠﻣﻮﻧ. ﺢﻴﺛ ﻙﺎﻧ ﺎﻟﺪﻠﻣﻮﻨﻳﻮﻧ ﻱﺪﻔﻧﻮﻧ ﻡﻮﺗﺎﻬﻣ ﻒﻳ ﺖﻠﻛ ﺎﻠﺗﻼﻟ ﻮﻬﻳ (ﻖﺑﻭﺭ) ﻮﻳﺪﻔﻧﻮﻧ ﻢﻌﻬﻣ ﺡﺎﺠﻳﺎﺘﻬﻣ ﺎﻠﻨﻔﻴﺳﺓ ﻭﺎﻠﺜﻤﻴﻧﺓ، ﻚﻣﺍ ﻱﺪﻔﻧﻮﻧ ﻢﻌﻬﻣ ﺞِـِﺭﺎﺑ ﺎﻠﻣﺍﺀ ﺎﻠﻔﺧﺍﺮﻳﺓ، ﻭﺎﻠﻤﻘﺘﻨﻳﺎﺗ ﺎﻠﺜﻤﻴﻧﺓ ﻆﻧًﺍ ﻢﻨﻬﻣ ﻭﺎﻌﺘﻗﺍﺩﺍ ﺄﻧ ﻩﺫﺍ ﺎﻠﻤﻴﺗ ﻕﺩ ﻲﻋﻭﺩ ﻞﻠﺤﻳﺍﺓ ﻒﻳ ﺄﻳ ﻞﺤﻇﺓ. [ﺐﺣﺎﺟﺓ ﻞﻤﺻﺩﺭ]

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 29, 2018, 10:05pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253/4 "2018-01-29T22:05:25Z")

</div>

Thanks, and nice to hear that it's working!

I'm trying to think of a better way to handle this from the user's perspective, but I'm not sure what the best solution is. We could introduce a `"writing_direction"` config setting, which might be more intuitive to write and less abstract than `"card_css": "direction: rtl"`... But it'd still be adding more complexity that's maybe not necessary. But I guess there are also too many RTL languages to have the front-end or back-end detect this automatically based on `nlp.lang`... and it'd also be making too many assumptions on the user's behalf. (And it still wouldn't solve the problem if you're using an `xx` model.) Anyway, if you have any ideas or suggestions, let me know!

> [@andy](#):
>
> There’s a separate tokenization problem, which is that spaCy is including periods in with the previous token in the xx and en models. But that’s an Arabic in spaCy issue, not a Prodigy issue.

Yes, this is likely because the [unicode character classes](https://github.com/explosion/spaCy/blob/master/spacy/lang/char_classes.py) used in the punctuation rules don't include the unicode ranges for Arabic. We actually just got two pull requests ([#1879](https://github.com/explosion/spaCy/pull/1879), [#1893](https://github.com/explosion/spaCy/pull/1893)) adding Persian to spaCy, and one of the PRs adds `\p{Arabic}` to the list of uncased characters. So this should probably solve the underlying issue.

---

<div class="post-metadata">

**Author:** ![andy](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/andy/32/966_2.png) [@andy](https://support.prodi.gy/u/andy)\
**Post date:** [January 29, 2018, 10:16pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253/5 "2018-01-29T22:16:11Z")

</div>

I think just having to specify `"card_css": "direction: rtl"` is just fine. That’s a pretty easy note in the README. I think people working with RTL languages in NLP are used to the RTL terminology.

I just checkout the PRs and hopefully that’ll fix it!

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 29, 2018, 11:00pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253/6 "2018-01-29T23:00:31Z")

</div>

Yeah, now that I think about it, maybe some type of auto-detection isn’t so bad after all. I mean, if the user isn’t aware of the config setting and just comes across this issue in `ner.manual`, it’s pretty difficult to figure out what’s going on. And we also don’t want people to think we’ve only ever thought about LTR languages, and everything else requires a “hack”.

So one alternative idea would be this: If all relevant recipes expose `'lang': nlp.lang` in their config, the controller could perform a simple check for the most common languages, and set the writing direction accordingly.

```python
if config.get('lang') in ('ar', 'fa', 'ur', 'he', 'yi'):
    config.setdefault('writing_direction', 'rtl')

```

The direction would only be set if no `"writing_direction"` setting is found in the existing user config. This would allow the user to manually set `"writing_direction": "ltr"` if their model is set to `"lang": "ar"`, but is actually an Arabic transliteration model.

---

<div class="post-metadata">

**Author:** ![andy](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/andy/32/966_2.png) [@andy](https://support.prodi.gy/u/andy)\
**Post date:** [January 29, 2018, 11:27pm UTC](https://support.prodi.gy/t/arabic-text-rendering-in-ner-manual/253/7 "2018-01-29T23:27:06Z")

</div>

I like that idea. There aren’t that many RTL languages in spaCy so it’s not bad to enumerate them. If it does set `writing_direction` to `rtl` based on the language, it might be nice to have a big message pop out in the log that tells users it’s auto switching and that they can override with `"writing_direction": "ltr"` in `.progidy.jsonl`.

The only recipe that wouldn’t really get covered with this is `mark`, where people will have text but no model to detect the language from. Maybe a hack that looks for characters in those languages and prints/logs a suggestion to set RTL? But that’s pretty clunky.
