# Annotated Data output formatting

**URL:** <https://support.prodi.gy/t/annotated-data-output-formatting/1211>\
**Category:** Uncategorized\
**Tags:** usage\
**Created:** [February 12, 2019, 10:18am UTC](https://support.prodi.gy/t/annotated-data-output-formatting/1211 "2019-02-12T10:18:29Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shhariar096](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/shhariar096/32/619_2.png) [@Shhariar096](https://support.prodi.gy/u/Shhariar096)\
**Post date:** [February 12, 2019, 10:18am UTC](https://support.prodi.gy/t/annotated-data-output-formatting/1211/1 "2019-02-12T10:18:29Z")

</div>

I am new to prodigy and this might be a silly question.  
i want to annotate PDF/doc/docx CV based on few labels. As prodigy does not support these formats directly, I converted cv data into a jsonl file and used **prodigy ner.manual dataset\_name en\_core\_web\_sm data.jsonl --label label1,label2,label3** to start prodigy app on localhost.  
After annotation i saved the output and retrieved it through `db-out`

here is the output file structure.  
`"text": "sample text....",`  
`"_input_hash":000000000,`  
`"_task_hash":0000000000,`  
`"tokens": [{ "text": "abc", "start": 0, "end": 2, "id": 0 },....],`  
`"spans": [{ "start": 0, "end": 16, "token_start": 0, "token_end": 1, "label": "xyz" }...], "answer": "accept"`

Now i want span’s content in this format:

```python
{"label":["label_name"],"points":[{"start":00,"end":50,"text":"abc xyz"}]}

```

**is there any mechanism by which i can map multiple tokens content into my label text by using start to end position of the text?** Here i am talking about multiple tokens because prodigy detects space separated characters as a token.  
suppose, **Albert Einstein was a theoretical physicist.** is a sentence. My annotator selects Albert Einstein as Name. now i want the output like this:

```python
{"label":["name"],"points":[{"start":00,"end":14,"text":"Albert Einstein"}]}

```

here **Albert** and **Einstein** are 2 tokens with token\_id 1 & 2.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [February 12, 2019, 11:57am UTC](https://support.prodi.gy/t/annotated-data-output-formatting/1211/2 "2019-02-12T11:57:51Z")

</div>

Hi! I hope I understand your question correctly – but what you describe is pretty much exactly what’s stored in the `"spans"` of your annotated data?

`"start"` and `"end"` are the character offsets into the text, `"token_start"` and `"token_end"` are the indices of the start and end tokens (corresponding with the tokens in `"tokens"`) and `"label"` is the label assigned to that particular span.
