# No Task Available

**URL:** <https://support.prodi.gy/t/no-task-available/4289>\
**Category:** Uncategorized\
**Tags:** ner, spacy, solved\
**Created:** [June 1, 2021, 3:30pm UTC](https://support.prodi.gy/t/no-task-available/4289 "2021-06-01T15:30:16Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 1, 2021, 3:30pm UTC](https://support.prodi.gy/t/no-task-available/4289/1 "2021-06-01T15:30:16Z")

</div>

Prodigy error "No Task Available" specific error "ERROR: Can't fetch tasks. Make sure the server is running correctly.

Can someone please help me in figuring this out this is a big data set but I was only able to get to 284, before I got this error message.

This is the code being run  
"# Prodigy using jsonl data that was converted from json in convert\_data.ipynb  
!python -m prodigy ner.manual test\_data\_5K blank:en ./test\_data\_5K.jsonl --label PER,ORG,MISC,LOC"

This is the output error messages

"Using 4 label(s): PER, ORG, MISC, LOC

✨ Starting the web server at [http://0.0.0.0:8088](http://0.0.0.0:8088) ...  
Open the app in your browser and start annotating!

Task exception was never retrieved  
future: \<Task finished name='Task-11' coro=\<RequestResponseCycle.run\_asgi() done, defined at /usr/local/anaconda3/lib/python3.8/site-packages/uvicorn/protocols/http/httptools\_impl.py:388\> exception=ValueError("Mismatched tokenization. Can't resolve span to token index 1. This can happen if your data contains pre-set spans. Make sure that the spans match spaCy's tokenization or add a 'tokens' property to your task.\n\n{'start': 1, 'end': 4, 'label': 'ORG'}")\>  
Traceback (most recent call last):  
File "/usr/local/anaconda3/lib/python3.8/site-packages/uvicorn/protocols/http/httptools\_impl.py", line 393, in run\_asgi  
self.logger.error(msg, exc\_info=exc)  
File "/usr/local/anaconda3/lib/python3.8/logging/ **init**.py", line 1463, in error  
self.\_log(ERROR, msg, args, \*\*kwargs)  
File "/usr/local/anaconda3/lib/python3.8/logging/ **init**.py", line 1577, in \_log  
self.handle(record)  
File "/usr/local/anaconda3/lib/python3.8/logging/ **init**.py", line 1586, in handle  
if (not self.disabled) and self.filter(record):  
File "/usr/local/anaconda3/lib/python3.8/logging/ **init**.py", line 807, in filter  
result = f.filter(record)  
File "cython\_src/prodigy/util.pyx", line 121, in prodigy.util.ServerErrorFilter.filter  
File "/usr/local/anaconda3/lib/python3.8/site-packages/uvicorn/protocols/http/httptools\_impl.py", line 390, in run\_asgi  
result = await app(self.scope, self.receive, self.send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/uvicorn/middleware/proxy\_headers.py", line 45, in **call**  
return await self.app(scope, receive, send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/fastapi/applications.py", line 140, in **call**  
await super(). **call** (scope, receive, send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/applications.py", line 134, in **call**  
await self.error\_middleware(scope, receive, send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/middleware/errors.py", line 178, in **call**  
raise exc from None  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/middleware/errors.py", line 156, in **call**  
await self.app(scope, receive, \_send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/middleware/cors.py", line 84, in **call**  
await self.simple\_response(scope, receive, send, request\_headers=headers)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/middleware/cors.py", line 140, in simple\_response  
await self.app(scope, receive, send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/middleware/base.py", line 25, in **call**  
response = await self.dispatch\_func(request, self.call\_next)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/prodigy/app.py", line 198, in reset\_db\_middleware  
response = await call\_next(request)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/middleware/base.py", line 45, in call\_next  
task.result()  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/middleware/base.py", line 38, in coro  
await self.app(scope, receive, send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/exceptions.py", line 73, in **call**  
raise exc from None  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/exceptions.py", line 62, in **call**  
await self.app(scope, receive, sender)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/routing.py", line 590, in **call**  
await route(scope, receive, send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/routing.py", line 208, in **call**  
await self.app(scope, receive, send)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/routing.py", line 41, in app  
response = await func(request)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/fastapi/routing.py", line 129, in app  
raw\_response = await run\_in\_threadpool(dependant.call, \*\*values)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/starlette/concurrency.py", line 25, in run\_in\_threadpool  
return await loop.run\_in\_executor(None, func, \*args)  
File "/usr/local/anaconda3/lib/python3.8/concurrent/futures/thread.py", line 57, in run  
result = self.fn(\*self.args, \*\*self.kwargs)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/prodigy/app.py", line 420, in get\_session\_questions  
return \_shared\_get\_questions(req.session\_id, excludes=req.excludes)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/prodigy/app.py", line 391, in \_shared\_get\_questions  
tasks = controller.get\_questions(session\_id=session\_id, excludes=excludes)  
File "cython\_src/prodigy/core.pyx", line 223, in prodigy.core.Controller.get\_questions  
File "cython\_src/prodigy/core.pyx", line 227, in prodigy.core.Controller.get\_questions  
File "cython\_src/prodigy/components/feeds.pyx", line 99, in prodigy.components.feeds.SharedFeed.get\_questions  
File "cython\_src/prodigy/components/feeds.pyx", line 106, in prodigy.components.feeds.SharedFeed.get\_next\_batch  
File "cython\_src/prodigy/components/feeds.pyx", line 245, in prodigy.components.feeds.RepeatingFeed.get\_session\_stream  
File "/usr/local/anaconda3/lib/python3.8/site-packages/toolz/itertoolz.py", line 376, in first  
return next(iter(seq))  
File "cython\_src/prodigy/components/preprocess.pyx", line 130, in add\_tokens  
File "cython\_src/prodigy/components/preprocess.pyx", line 222, in prodigy.components.preprocess.\_add\_tokens  
File "cython\_src/prodigy/components/preprocess.pyx", line 199, in prodigy.components.preprocess.sync\_spans\_to\_tokens  
ValueError: Mismatched tokenization. Can't resolve span to token index 1. This can happen if your data contains pre-set spans. Make sure that the spans match spaCy's tokenization or add a 'tokens' property to your task.

{'start': 1, 'end': 4, 'label': 'ORG'}"

---

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 1, 2021, 6:15pm UTC](https://support.prodi.gy/t/no-task-available/4289/2 "2021-06-01T18:15:10Z")

</div>

Can someone assist me with this issue?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 2, 2021, 2:55am UTC](https://support.prodi.gy/t/no-task-available/4289/3 "2021-06-02T02:55:50Z")

</div>

Hi! It's usually not so helpful to bump a thread, especially not after such a short period of time – it can often make it harder for us to keep track of new threads and make sure we can answer everyone.

If you see an error message in the UI about the server not running correctly, it typically means that the Python recipe raised an error. In this case, the cause of the error was this:

> [@jtsolomon](#):
>
> ValueError: Mismatched tokenization. Can't resolve span to token index 1. This can happen if your data contains pre-set spans. Make sure that the spans match spaCy's tokenization or add a 'tokens' property to your task.
> 
> {'start': 1, 'end': 4, 'label': 'ORG'}"

Which recipe are you running and how are you loading in the data? Are you labelling examples with pre-defined spans?

Basically, what the error means is that Prodigy came across an annotated span referring to a slice from character 1 to 4 – but this slice doesn't map to valid tokens produced by the tokenizer. If you've generated pre-annotated spans yourself, this can sometimes happen if you have leading or trailing whitespace, or an off-by-one error in the offsets. In some cases, it can also mean that your data assumes that a string is split when it isn't – for instance, `A-B` → `["A", "-", "B"]` vs. `["A-B"]`. So you can either adjust the tokenization rules, provide your custom `"tokens"` in the data (if you'll be using the same tokenization during training) or adjust your annotation scheme to ensure that the model will be able to learn from the data. Also see this thread for tips on how to find and resolve mismatched tokenization:

> [@Matching tokenisation on pre-existing annotated data](https://support.prodi.gy/t/matching-tokenisation-on-pre-existing-annotated-data/2710):
>
> We have pre-existing data annotated for NER that I'd like to use prodigy to review and correct before training an NER model on it. It's in the (text, {entitiies: ...}) spacy training format but not tokenised. Having converted this to {text: ..., spans: ...} as input to ner.manual, a few examples load fine but some of the existing spans do not exactly land on the tokens the model is generating and I see a 'ValueError: Mismatched tokenization'. We'd like to 'snap' the existing spans to agree wit…

---

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 2, 2021, 3:03am UTC](https://support.prodi.gy/t/no-task-available/4289/4 "2021-06-02T03:03:54Z")

</div>

HI, Thank you for replying back.

this is the data that I have been using. after line 284 that is when I got this error message. Some sentience's have gotten labeled already and I want to check if it was done correctly.  
(it might be easier to talk over the phone to fix this issue, is there a phone number to call?)

As you can see the format of the data is all the same, hence my confusion why it stoppered after line 284. and saying error "ValueError: Mismatched tokenization. Can't resolve span to token index 1. This can happen if your data contains pre-set spans. Make sure that the spans match spaCy's tokenization or add a 'tokens' property to your task.

{'start': 1, 'end': 4, 'label': 'ORG'}" Clearly you can see that the tokens are the same throughout and using the same format.

This is my run command for Prodigy  
!python -m prodigy ner.manual test\_data\_5K blank:en ./test\_data\_5K.jsonl --label PER,ORG,MISC,LOC

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/2/2a7a82fd2b24a31ffe48c88fea4f3a940ed57333.png)

I looked at your message, and took a look at the text for the first line not sure why but i think """" might be the issue. when I deleted them from the first line of data (see below) I get another error message (see below)

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/4/492620768c8d38b311761b2a006ac4db5e810b5f.png)

Using 4 label(s): PER, ORG, MISC, LOC  
Traceback (most recent call last):  
File "/usr/local/anaconda3/lib/python3.8/runpy.py", line 194, in \_run\_module\_as\_main  
return \_run\_code(code, main\_globals, None,  
File "/usr/local/anaconda3/lib/python3.8/runpy.py", line 87, in \_run\_code  
exec(code, run\_globals)  
File "/usr/local/anaconda3/lib/python3.8/site-packages/prodigy/ **main**.py", line 53, in   
controller = recipe(\*args, use\_plac=True)  
File "cython\_src/prodigy/core.pyx", line 331, in prodigy.core.recipe.recipe\_decorator.recipe\_proxy  
File "cython\_src/prodigy/core.pyx", line 353, in prodigy.core.\_components\_to\_ctrl  
File "cython\_src/prodigy/core.pyx", line 142, in prodigy.core.Controller. **init**  
File "cython\_src/prodigy/components/feeds.pyx", line 56, in prodigy.components.feeds.SharedFeed. **init**  
File "cython\_src/prodigy/components/feeds.pyx", line 155, in prodigy.components.feeds.SharedFeed.validate\_stream  
File "/usr/local/anaconda3/lib/python3.8/site-packages/toolz/itertoolz.py", line 376, in first  
return next(iter(seq))  
File "cython\_src/prodigy/components/preprocess.pyx", line 130, in add\_tokens  
File "cython\_src/prodigy/components/preprocess.pyx", line 222, in prodigy.components.preprocess.\_add\_tokens  
File "cython\_src/prodigy/components/preprocess.pyx", line 199, in prodigy.components.preprocess.sync\_spans\_to\_tokens  
ValueError: Mismatched tokenization. Can't resolve span to token index 101. This can happen if your data contains pre-set spans. Make sure that the spans match spaCy's tokenization or add a 'tokens' property to your task.

{'start': 101, 'end': 112, 'label': 'PER'}

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [June 2, 2021, 9:19pm UTC](https://support.prodi.gy/t/no-task-available/4289/5 "2021-06-02T21:19:48Z")

</div>

Hi!

As Ines explained, some of your input data has an issue with tokenization not matching the tokenizer used during annotation of the model. For instance, if your tokenizer says that `"Tim's house"` are two words, `"Tim's"` and `"house"`, then you won't be able to annotate just `"Tim"` as a `Span` because that wouldn't align to token boundaries. So for that example, you'd either have to change your span annotation, or your tokenizer, or provide the token annotations in your input file.

I can't `grep` through your screenshot of the data, but you must have a line somewhere with a span `{'start': 1, 'end': 4, 'label': 'ORG'}"` - this particular line will likely be the culprit.

---

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 2, 2021, 9:23pm UTC](https://support.prodi.gy/t/no-task-available/4289/6 "2021-06-02T21:23:23Z")

</div>

Are you saying that a sentence with the name Bill Gates can not be grouped together?  
I am assuming and using Prodigy by highlighting the entire name and saying this is a Person. What you seem to suggest I have to say Bill is a person and Gates is a person and I can not say Bill Gates is a person.

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [June 2, 2021, 9:30pm UTC](https://support.prodi.gy/t/no-task-available/4289/7 "2021-06-02T21:30:03Z")

</div>

No, that's not what I was trying to say.

The first step in any NLP pipeline is called "tokenization" - it means breaking up a sentence into tokens or words. Typically, a sentence "Bill Gates is rich" will be broken up into 4 tokens: ["Bill", "Gates", "is", "rich"].

Tokenization can be slightly different depending on the exact rules that are being used though. For instance, the sentence "I won't go". Can be parsed as 3 tokens, keeping "won't" together as one token, or as 4 tokens, splitting up "won't". It depends on the tokenizer.

When you're annotatings spans with Prodigy, a Span always needs to follow those token boundaries. That means that a span needs to include full tokens, so it can be "Bill Gates" - spanning two tokens. However, you typically can't annotate the substring "Gate" in "Gates" if "Gates" is one token.

That is what your error message is about: you have a span annotation that does not align with the token boundaries defined by the tokenizer you're using in your recipe. The error message helps you to identify the "offending" input. If you can cite it here, we can likely help you fix that particular example.

---

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 2, 2021, 9:54pm UTC](https://support.prodi.gy/t/no-task-available/4289/8 "2021-06-02T21:54:38Z")

</div>

Okay I understand what you are saying.

original error message is this, below, I am assuming that line one in the jsonl file is the issue as that is what the token index is saying.

ValueError: Mismatched tokenization. Can't resolve span to token index 1. This can happen if your data contains pre-set spans. Make sure that the spans match spaCy's tokenization or add a 'tokens' property to your task.

{'start': 1, 'end': 4, 'label': 'ORG'}

This is the first line of the file

 ![image](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/3/32ad18889190f581970090d5d6515f927392d883.png)

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [June 3, 2021, 7:11am UTC](https://support.prodi.gy/t/no-task-available/4289/9 "2021-06-03T07:11:44Z")

</div>

No, the error is not saying that there is a problem with the first line.

> Can't resolve span to token index 1.  
> {'start': 1, 'end': 4, 'label': 'ORG'}

There is a particular line with an ORG entity annotated from char 1 to 4, which doesn't align to token boundaries. The span can't start at index 1, because that's not a token start.

Can you find the data line with ` {'start': 1, 'end': 4, 'label': 'ORG'}` in it?

---

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 3, 2021, 5:36pm UTC](https://support.prodi.gy/t/no-task-available/4289/10 "2021-06-03T17:36:26Z")

</div>

Hi,

I looked at the data and there is no {'start': 1, 'end': 4, 'label': 'ORG'} present.  
Not sure why this is the error message as there is no span that uses start 1 and end 4 for ORG label.

please assist. I can attach the data file if you would like but i am not sure how to attach it.

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [June 4, 2021, 7:44am UTC](https://support.prodi.gy/t/no-task-available/4289/11 "2021-06-04T07:44:40Z")

</div>

When you hit reply, there should be an upload icon that let's you attach files to your post. If you attach your dataset (not a screenshot), I can have a look 😉

---

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 7, 2021, 5:03pm UTC](https://support.prodi.gy/t/no-task-available/4289/12 "2021-06-07T17:03:30Z")

</div>

here is the JSONL file, Thank you.  
[test\_data\_5K.jsonl](https://support.prodi.gy/uploads/short-url/reh2zm0kfKEx6fsdGrpRFfy94NW.jsonl) (1.0 MB)

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 8, 2021, 2:24am UTC](https://support.prodi.gy/t/no-task-available/4289/13 "2021-06-08T02:24:58Z")

</div>

Here's the example that defines the span with `{'start': 1, 'end': 4, 'label': 'ORG'}`:

```json
{"text":"/LIC traditionally has to push through layers of bureaucracy to get anything done.","spans":[{"start":1,"end":4,"label":"ORG"}]}

```

As @SofieVL explained earlier, the problem here is that `/LIC` is treated as one token, whereas the span expects to describe a token starting at character 1 – but that will never be true using the tokenization provided. This means that the example annotates token-based tags that will never match, since the tokens they describe aren't actually produced. Prodigy alerts you to that because it means that otherwise, you'd be creating annotations that your model can't learn from.

There are a few similar instances in the data that you can find by searching for `"start": 1,`, but I think they're mostly texts starting with quotes, which the tokenizer would split by default. But it might still be worth double-checking them. (The thread I linked above includes code you can use to quickly check your annotated spans against a given tokenization and find mismatches programmatically.)

If the mismatches are caused by the data not being "clean", e.g. leftover markup etc., adding a preprocessing step can help – just make sure the same preprocessing is also applied at runtime later so the model doesn't see anything unexpected. Alternatively, if arbitrary trailing characters are common in your data, another solution would be to add some stricter tokenization rules that split off characters like `/` at the beginning of a word, so you end up with `["/", "LIC"]` instead of `["/LIC"]`.

---

<div class="post-metadata">

**Author:** ![jtsolomon](https://avatars.discourse-cdn.com/v4/letter/j/58956e/32.png) [@jtsolomon](https://support.prodi.gy/u/jtsolomon)\
**Post date:** [June 9, 2021, 6:25pm UTC](https://support.prodi.gy/t/no-task-available/4289/14 "2021-06-09T18:25:57Z")

</div>

Thank you for your help in fining the issue the issue was for "/" character in between two words without a space in between. I corrected this and added a space. I found another error after this that pointed to a "'" next to a word.

I was converting a JSON to a JSONL file, because Prodigy needs a JSONl file type. I tried to fix the spacing issues from having stand alone tokens to a sentence format. I suppose I need to make a better rules in deciding how to concatenate the tokens into a nice string with correct spacings. It seems my spacing rules that I made was the cause of this.

That said in an earlier reply from one of your staff you indicated that we can customize how it looks and parses for tokens. can you explain to me how to customize this further. That maybe a route to take to have clean data that Prodigy can do NER on.

---

<div class="post-metadata">

**Author:** ![SofieVL](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/sofievl/32/915_2.png) [@SofieVL](https://support.prodi.gy/u/SofieVL)\
**Post date:** [June 10, 2021, 11:48am UTC](https://support.prodi.gy/t/no-task-available/4289/15 "2021-06-10T11:48:46Z")

</div>

Have you looked at this thread that Ines linked earlier for dealing with non-matching tokens? [Matching tokenisation on pre-existing annotated data](https://support.prodi.gy/t/matching-tokenisation-on-pre-existing-annotated-data/2710)

The other option is indeed to define the `tokens` in your input data - then the tokenizer won't run and will just take the tokens as you've defined them. You can see the expected format here: [https://prodi.gy/docs/api-interfaces#ner\_manual](https://prodi.gy/docs/api-interfaces#ner_manual)

```python
{
  "text": "First look at the new MacBook Pro",
  "spans": [
    {"start": 22, "end": 33, "label": "PRODUCT", "token_start": 5, "token_end": 6}
  ],
  "tokens": [
    {"text": "First", "start": 0, "end": 5, "id": 0},
    {"text": "look", "start": 6, "end": 10, "id": 1},
    {"text": "at", "start": 11, "end": 13, "id": 2},
    {"text": "the", "start": 14, "end": 17, "id": 3},
    {"text": "new", "start": 18, "end": 21, "id": 4},
    {"text": "MacBook", "start": 22, "end": 29, "id": 5},
    {"text": "Pro", "start": 30, "end": 33, "id": 6}
  ]
}

```

(you don't need to have the `spans` predefined. If not, you have to annotate them all yourself)
