# Testing add\_tokens

**URL:** <https://support.prodi.gy/t/testing-add-tokens/628>\
**Category:** Uncategorized\
**Tags:** usage, solved\
**Created:** [June 19, 2018, 7:55pm UTC](https://support.prodi.gy/t/testing-add-tokens/628 "2018-06-19T19:55:48Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![pvcastro](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/pvcastro/32/276_2.png) [@pvcastro](https://support.prodi.gy/u/pvcastro)\
**Post date:** [June 19, 2018, 7:55pm UTC](https://support.prodi.gy/t/testing-add-tokens/628/1 "2018-06-19T19:55:48Z")

</div>

Hi there,

I was running some tests trying to fix a Mismatched tokenization error I’m getting, and I tried running the following code, that was recommended in another thread:

```
from prodigy.components.preprocess import add_tokens
import en_core_web_sm

nlp = en_core_web_sm.load()
text = " The upstart streaming service, which is primarily geared for sports fans, has an uphill climb against deep-pocketed competitors marketing cable alternatives to cord-cutters: YouTube TV, Hulu Live and Sony's PlayStation Vue."
stream = [{'text': text, 'spans': {'start': 175, 'end': 185}}]
new_stream = add_tokens(nlp, stream)
print(list(new_stream))

```

I’m getting the following exception when running this code:

```
TypeError Traceback (most recent call last)
<ipython-input-115-958a6dcd96e1> in <module>()
  6 stream = [{'text': text, 'spans': {'start': 175, 'end': 185}}]
  7 new_stream = add_tokens(nlp, stream)
----> 8 print(list(new_stream))

cython_src/prodigy/components/preprocess.pyx in add_tokens()

TypeError: string indices must be integers

```

Thanks!

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 20, 2018, 7:46am UTC](https://support.prodi.gy/t/testing-add-tokens/628/2 "2018-06-20T07:46:44Z")

</div>

I think you might actually have a small typo in your stream: `"spans"` here is a dictionary, when it should be a list of dictionaries.

The new validation mechanism should catch errors like that within the recipes. If you want to implement something like this yourself (e.g. to make sure that your stream is formatted correctly for the annotation task), you can also call into the validator directly.

```python
from prodigy.components.validate import Validator

validator = Validator('ner_manual') # the view_id you want to use the stream with
for eg in stream:
    validator.check(eg)

```

**Disclaimer:** This is currently internals only, so the API may change in the future.

---

<div class="post-metadata">

**Author:** ![pvcastro](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/pvcastro/32/276_2.png) [@pvcastro](https://support.prodi.gy/u/pvcastro)\
**Post date:** [June 21, 2018, 2:11pm UTC](https://support.prodi.gy/t/testing-add-tokens/628/3 "2018-06-21T14:11:55Z")

</div>

Thanks!

I copied the code exactly from [here](https://support.prodi.gy/t/valueerror-mismatched-tokenization-in-ner-make-gold/292), but adding the brackets solved it!

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 21, 2018, 2:31pm UTC](https://support.prodi.gy/t/testing-add-tokens/628/4 "2018-06-21T14:31:52Z")

</div>

Ah, sorry, that was also a typo in my semi-pseudocode then! Thanks, fixed 👍
