# empty spans and spans with no 'text' attribute

**URL:** <https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224>\
**Category:** Uncategorized\
**Tags:** database, solved\
**Created:** [January 10, 2023, 2:32am UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224 "2023-01-10T02:32:07Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![stefan.bartell](https://avatars.discourse-cdn.com/v4/letter/s/3d9bf3/32.png) [@stefan.bartell](https://support.prodi.gy/u/stefan.bartell)\
**Post date:** [January 10, 2023, 2:32am UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/1 "2023-01-10T02:32:07Z")

</div>

Hi. I'm having two issues. The output of db-out when I save annotations produces empty spans and spans with no 'text' attribute, while when I used db-out before, spans were an empty list ([]) or a list of dictionaries with a 'text' attribute (key). I need the 'text' attribute to know what the span refers to in the input text. Here is a minimal working example where I try to recreate the problem.

First, here is the command I use to start annotations:

```python
python3 -m prodigy ner.manual identify_dosage_non_dosage_validate_data_SB2 en_core_web_lg validate_data_dosage_annotations_SB2.jsonl --label non_dosage,dosage

```

which calls the file:  
[validate\_data\_dosage\_annotations\_SB2.jsonl](https://support.prodi.gy/uploads/short-url/pkmt91xp4ufYby1q0cdpDd7ki62.jsonl) (52.3 KB)

Next, I perform the annotations, annotating text as a dosage or non-dosage.  
Then I save the annotations with

```python
python3 -m prodigy db-out identify_dosage_non_dosage_validate_data_SB2 > validate_data_dosage_non_dosage_annotations_SB2.jsonl

```

The annotations are saved here:  
[validate\_data\_dosage\_non\_dosage\_annotations\_SB2.jsonl](https://support.prodi.gy/uploads/short-url/uD71N0hzQGqTSmfX93TrwWzFnhL.jsonl) (53.1 KB)

Now, in order to visualize the annotations in a spreadsheet, I use this script in Python:

```python
import pandas as pd

df_jsonl_annotations = pd.read_json('validate_data_dosage_non_dosage_annotations_SB2.jsonl', lines=True)

df_jsonl_annotations.to_csv('validate_data_dosage_non_dosage_annotations_SB2.csv', index=False)

```

The results can be seen here (originally a CSV file) (which I have truncated for readability):

```python
text	_input_hash	_task_hash	_is_binary	tokens	_view_id	answer	_timestamp	spans
January 9 - 241 6 - 375mg split into 3 doses. 96m deadlift/back/shoulder session. 30m cardio. 7,872 steps. 1,640 calories at 17g (7g net) carbs, 93g fat, 128g protein. 1 5g water January 10 - 241 4 - 375mg	1201376478	-478339982	FALSE	[{'text': 'January', 'start': 0, 'end': 7, 'id': 0, 'ws': True}, {'text': '9', 'start': 8, 'end': 9, 'id': 1, 'ws': True}, {'text': '-', 'start': 10, 'end': 11, 'id': 2, 'ws': True}	ner_manual	accept	1673316381	[{'start': 18, 'end': 25, 'token_start': 5, 'token_end': 7, 'label': 'non_dosage'}, {'start': 125, 'end': 143, 'token_start': 32, 'token_end': 39, 'label': 'non_dosage'}
 Originally Posted by itismethebeeFirst off, I turned 18 this year.	-681172102	1046073490	FALSE	[{'text': ' ', 'start': 0, 'end': 1, 'id': 0, 'ws': False}, {'text': 'Originally', 'start': 1, 'end': 11, 'id': 1, 'ws': True}, {'text': 'Posted', 'start': 12, 'end': 18, 'id': 2, 'ws': True}				
d': 484 'ws': True} {'text': 'this' 'start': 2191 'end': 2195 'id': 485 'ws': True} {'text': 'post'	

```

As you can see, the first line contains spans with no 'text' attribute, while the second contains empty spans, which I don't think should be the case.

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [January 10, 2023, 7:56pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/2 "2023-01-10T19:56:03Z")

</div>

hi @stefan.bartell!

Thanks for your question and welcome to the Prodigy community 👋

> [@stefan.bartell](#):
>
> The output of db-out when I save annotations produces empty spans and spans with no 'text' attribute, while when I used db-out before, spans were an empty list () or a list of dictionaries with a 'text' attribute (key).

This is a bit odd. Your annotated file has `"answer":"accept"` tags for each record, indicating they were accepted (saved), but yes, they should include your `spans` as a list of dictionaries, dictionary per span. For `ner.manual`, saved annotated spans will be in `spans`. See [this link](https://prodi.gy/docs/api-interfaces#ner_manual) for what the data looks like for it.

When I did the steps below, everything worked out fine:

```python
python3 -m prodigy ner.manual identify_dosage_non_dosage_validate_data_SB2 en_core_web_lg validate_data_dosage_annotations_SB2.jsonl --label non_dosage,dosage

```

Then annotated two spans, accepting them by clicking the Green "Accept" button.

 ![localhost_8080_ (16)](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/a/ade262d408d46d05e33ed160f6bbbf78d5075858.png)

Then on the next record, I clicked save at the top:

 ![localhost_8080_ (17)](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/6/6dddea354157b1e2a00aab08150d085f633011c7.png)

I can now go back to my terminal and shut down the server by pressing CTRL + C. When I do this, you can also confirm whether your annotation was saved in the CLI:

```python
$ python3 -m prodigy ner.manual dosage_dataset en_core_web_lg data/validate_data_dosage_annotations_SB2.jsonl --label non_dosage,dosage
Using 2 label(s): non_dosage, dosage
Added dataset dosage_dataset to database SQLite.

✨ Starting the web server at http://localhost:8080 ...
Open the app in your browser and start annotating!

^C
✔ Saved 1 annotations to database SQLite
Dataset: dosage_dataset
Session ID: 2023-01-10_14-17-06

```

Now if I output out that file with `db-out`:

```python
python3 -m prodigy db-out dosage_dataset > dosage.jsonl

```

I get:

```python
{
  "text": "January 9 - 241 6 - 375mg split into 3 doses. 96m deadlift/back/shoulder session. 30m cardio. 7,872 steps. 1,640 calories at 17g (7g net) carbs, 93g fat, 128g protein. 1 5g water January 10 - 241 4 - 375mg split into 3 doses. 74m arms session. 30m cardio. 8,402 steps. 1,640 calories at 26g (15g net) carbs, 127g fat, 106g protein. 1 5g water Expected a weight drop by now so I hope it's just water retention as I've been religious with everything and eating has been on point. My circadian rhythms do usually ebb and flow where I'll get a \"whoosh\" weight drop every once in awhile. I'm gonna keep on, keepin' on.",
  "_input_hash": 1201376478,
  "_task_hash": -478339982,
  "_is_binary": false,
  "tokens": [
    {
      "text": "January",
      "start": 0,
      "end": 7,
      "id": 0,
      "ws": true
    },
    .
    .
    .
    {
      "text": ".",
      "start": 612,
      "end": 613,
      "id": 163,
      "ws": false
    }
  ],
  "_view_id": "ner_manual",
  "answer": "accept",
  "_timestamp": 1673378743,
  "spans": [
    {
      "start": 12,
      "end": 44,
      "token_start": 3,
      "token_end": 11,
      "label": "dosage"
    },
    {
      "start": 192,
      "end": 224,
      "token_start": 56,
      "token_end": 64,
      "label": "dosage"
    }
  ]
}

```

These are the two spans that were saved.

Now those spans do not include by default the raw span text. You can add this by modifying your `db-out` recipe:

> [@spans.manual merge tokens using db-out](https://support.prodi.gy/t/spans-manual-merge-tokens-using-db-out/6121/2):
>
> Hi @cheyanneb! So you're just looking to add the raw span text to each span dict? Thinking you could just add this to db-out: for eg in examples: for span in eg["spans"]: span['text'] = eg['text'][span['start']:span['end']] If you create a new flag argument (add\_span\_text) to turn this on or off (set off by default), you could run this from pathlib import Path from typing import Optional, Union import srsly from prodigy.components.db import connect from prodigy.util import m…

Can you double check that you annotated correctly (e.g., highlighting the spans, clicking Accept (green) button, and saving your annotations by clicking the Save Button)?

---

<div class="post-metadata">

**Author:** ![stefan.bartell](https://avatars.discourse-cdn.com/v4/letter/s/3d9bf3/32.png) [@stefan.bartell](https://support.prodi.gy/u/stefan.bartell)\
**Post date:** [January 10, 2023, 8:56pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/3 "2023-01-10T20:56:19Z")

</div>

Thanks @ryanwesslen. It looks like you were able to reproduce the problem. I read your reply to @cheyanneb. The code you suggested to add to `db-out`:

```python
for eg in examples:
    for span in eg["spans"]:
          span['text'] = eg['text'][span['start']:span['end']]

```

You indicated that `db-out` is located in `commands.py`. What is the path to `commands.py`? I have multiple files with that name on my machine, but I can't tell if any of them are for Prodigy.

As for the other code you suggested:

```python
from pathlib import Path
from typing import Optional, Union

import srsly
from prodigy.components.db import connect
from prodigy.util import msg

def db_out(
    set_id: str,
    out_dir: Optional[Union[str, Path]] = None,
    answer: str = None,
    flagged_only: bool = False,
    dry: bool = False,
    add_span_text: bool = False,
) -> None:
    """
    Export annotations from the database. Files will be exported in
    Prodigy's JSONL format.
    """
    DB = connect()
    if set_id not in DB:
        msg.fail(f"Can't find '{set_id}' in database {DB.db_name}", exits=1)
    examples = DB.get_dataset_examples(set_id)
    if flagged_only:
        examples = [eg for eg in examples if eg.get("flagged")]
    if answer:
        examples = [eg for eg in examples if eg.get("answer") == answer]

    # add span text
    if add_span_text:
        for eg in examples:
            for span in eg["spans"]:
                span['text'] = eg['text'][span['start']:span['end']]

    if out_dir is None:
        for eg in examples:
            print(srsly.json_dumps(eg))
    else:
        out_dir = Path(out_dir)
        if not out_dir.exists():
            out_dir.mkdir()
        out_file = out_dir / f"{set_id}.jsonl"
        if not dry:
            srsly.write_jsonl(out_file, examples)
        msg.good(
            f"Exported {len(examples)} annotations from '{set_id}' in database {DB.db_name}",
            out_file.resolve(),
        )

```

I wasn't sure how to run this. Am I adding it to a file or is it in its own file? Is it in `my_dbout_script.py`, which is run with `-F my_dbout_script.py`? How do I set `add_span_text` to `True`?

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [January 10, 2023, 10:17pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/4 "2023-01-10T22:17:21Z")

</div>

> [@stefan.bartell](#):
>
> What is the path to `commands.py`?

This is based on where your Prodigy library is installed.

Type in `python -m prodigy stats` then find the `Location:` folder. From there, open that Location path in a window and look for the `recipes/commands.py` script. (FYI you can find other Prodigy recipes in that `recipes/` folder too).

> [@stefan.bartell](#):
>
> I wasn't sure how to run this. Am I adding it to a file or is it in its own file? Is it in `my_dbout_script.py`, which is run with `-F my_dbout_script.py`? How do I set `add_span_text` to `True`?

Yes, the easiest way would be to run it as a [custom recipe](https://prodi.gy/docs/custom-recipes#writing). But to run it from the command line, you will need to wrap the `@prodigy.recipe` decorator around your function.

```python
@prodigy.recipe(
    "db-out",
    set_id=("Name of dataset to export", "positional", None, str),
    out_dir=("Path to output directory", "positional", None, str),
    answer=("Only export annotations with this answer", "option", "a", str),
    flagged_only=("DEPRECATED: Only export flagged annotations", "flag", "F", bool),
    dry=("Perform a dry run", "flag", "D", bool),
    add_span_text=("Flag to add in the text spans", "flag", None, bool),
)

```

If you're new to decorators, here's a great [tutorial](https://calmcode.io/decorators/introduction.html) on them.

Then you should be able to run:

```python
python -m prodigy db-out my_dataset --add_span_text -F my_dbout_script.py

```

---

<div class="post-metadata">

**Author:** ![stefan.bartell](https://avatars.discourse-cdn.com/v4/letter/s/3d9bf3/32.png) [@stefan.bartell](https://support.prodi.gy/u/stefan.bartell)\
**Post date:** [January 10, 2023, 10:35pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/5 "2023-01-10T22:35:14Z")

</div>

Can you show the code snippet for where

```python
for eg in examples:
        for span in eg["spans"]:
          span['text'] = eg['text'][span['start']:span['end']]

```

is added to `commands.py`? I tried adding it to `@recipe( "db-out"` and `def db_out(` but got syntax errors.

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [January 11, 2023, 3:31pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/6 "2023-01-11T15:31:01Z")

</div>

> [@stefan.bartell](#):
>
> Can you show the code snippet for where
> 
> ```python
> for eg in examples:
> for span in eg["spans"]:
> span['text'] = eg['text'][span['start']:span['end']]
> 
> ```
> 
> is added to `commands.py`?

This code snippet isn't included in the `commands.py` by default. It's why I was suggesting that if you want to run something like `db-out`, likely your best bet is to create a custom recipe and run that separately. Let me know if you have any further questions!

---

<div class="post-metadata">

**Author:** ![stefan.bartell](https://avatars.discourse-cdn.com/v4/letter/s/3d9bf3/32.png) [@stefan.bartell](https://support.prodi.gy/u/stefan.bartell)\
**Post date:** [January 11, 2023, 9:02pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/7 "2023-01-11T21:02:08Z")

</div>

Hi @ryanwesslen. I appreciate your suggestions. Here is what I tried: saved the code here as `my_dbout_script.py`:

```python
import prodigy
from pathlib import Path
from typing import Optional, Union

import srsly
from prodigy.components.db import connect
from prodigy.util import msg

@prodigy.recipe(
    "db-out",
    set_id=("Name of dataset to export", "positional", None, str),
    out_dir=("Path to output directory", "positional", None, str),
    answer=("Only export annotations with this answer", "option", "a", str),
    flagged_only=("DEPRECATED: Only export flagged annotations", "flag", "F", bool),
    dry=("Perform a dry run", "flag", "D", bool),
    add_span_text=("Flag to add in the text spans", "flag", None, bool),
)

def db_out(
    set_id: str,
    out_dir: Optional[Union[str, Path]] = None,
    answer: str = None,
    flagged_only: bool = False,
    dry: bool = False,
    add_span_text: bool = False,
) -> None:
    """
    Export annotations from the database. Files will be exported in
    Prodigy's JSONL format.
    """
    DB = connect()
    if set_id not in DB:
        msg.fail(f"Can't find '{set_id}' in database {DB.db_name}", exits=1)
    examples = DB.get_dataset_examples(set_id)
    if flagged_only:
        examples = [eg for eg in examples if eg.get("flagged")]
    if answer:
        examples = [eg for eg in examples if eg.get("answer") == answer]

    # add span text
    if add_span_text:
        for eg in examples:
            for span in eg["spans"]:
                span['text'] = eg['text'][span['start']:span['end']]

    if out_dir is None:
        for eg in examples:
            print(srsly.json_dumps(eg))
    else:
        out_dir = Path(out_dir)
        if not out_dir.exists():
            out_dir.mkdir()
        out_file = out_dir / f"{set_id}.jsonl"
        if not dry:
            srsly.write_jsonl(out_file, examples)
        msg.good(
            f"Exported {len(examples)} annotations from '{set_id}' in database {DB.db_name}",
            out_file.resolve(),
        )

```

Then ran the command:

```python
python3 -m prodigy db-out identify_dosage_non_dosage_validate_data_SB2 > validate_data_dosage_non_dosage_annotations_SB2.jsonl -add_span_text -F my_dbout_script.py

```

But this resulted in a blank `validate_data_dosage_non_dosage_annotations_SB2.jsonl` file. I also reran the annotations to make sure they weren't empty. Do you have an idea of what I am doing wrong?

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [January 11, 2023, 9:27pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/8 "2023-01-11T21:27:37Z")

</div>

Thinking more, let's just eliminate the `-add_span_text` and make it by default it always adds the texts. If you don't want to include it, you can just use the standard `db-out`.

So use this script:

```python
#my_dbout_script
import prodigy
from pathlib import Path
from typing import Optional, Union

import srsly
from prodigy.components.db import connect
from prodigy.util import msg

@prodigy.recipe(
    "db-out",
    set_id=("Name of dataset to export", "positional", None, str),
    out_dir=("Path to output directory", "positional", None, str),
    answer=("Only export annotations with this answer", "option", "a", str),
    flagged_only=("DEPRECATED: Only export flagged annotations", "flag", "F", bool),
    dry=("Perform a dry run", "flag", "D", bool),
)

def db_out(
    set_id: str,
    out_dir: Optional[Union[str, Path]] = None,
    answer: str = None,
    flagged_only: bool = False,
    dry: bool = False,
) -> None:
    """
    Export annotations from the database. Files will be exported in
    Prodigy's JSONL format.
    """
    DB = connect()
    if set_id not in DB:
        msg.fail(f"Can't find '{set_id}' in database {DB.db_name}", exits=1)
    examples = DB.get_dataset_examples(set_id)
    if flagged_only:
        examples = [eg for eg in examples if eg.get("flagged")]
    if answer:
        examples = [eg for eg in examples if eg.get("answer") == answer]

    for eg in examples:
        if eg.get('spans') is not None:
            for span in eg.get('spans'):
                span['text'] = eg['text'][span['start']:span['end']]

    if out_dir is None:
        for eg in examples:
            print(srsly.json_dumps(eg))
    else:
        out_dir = Path(out_dir)
        if not out_dir.exists():
            out_dir.mkdir()
        out_file = out_dir / f"{set_id}.jsonl"
        if not dry:
            srsly.write_jsonl(out_file, examples)
        msg.good(
            f"Exported {len(examples)} annotations from '{set_id}' in database {DB.db_name}",
            out_file.resolve(),
        )

```

Then try this:

```python
python3 -m prodigy db-out identify_dosage_non_dosage_validate_data_SB2 -F my_dbout_script.py > validate_data_dosage_non_dosage_annotations_SB2.jsonl

```

Then I get the spans:

```python
{
  "text": "January 9 - 241 6 - 375mg split into 3 doses. 96m deadlift/back/shoulder session. 30m cardio. 7,872 steps. 1,640 calories at 17g (7g net) carbs, 93g fat, 128g protein. 1 5g water January 10 - 241 4 - 375mg split into 3 doses. 74m arms session. 30m cardio. 8,402 steps. 1,640 calories at 26g (15g net) carbs, 127g fat, 106g protein. 1 5g water Expected a weight drop by now so I hope it's just water retention as I've been religious with everything and eating has been on point. My circadian rhythms do usually ebb and flow where I'll get a \"whoosh\" weight drop every once in awhile. I'm gonna keep on, keepin' on.",
  "_input_hash": 1201376478,
  "_task_hash": -478339982,
  "_is_binary": false,
  "tokens": [
    {
      "text": "January",
      "start": 0,
      "end": 7,
      "id": 0,
      "ws": true
    },
    {
      "text": "9",
      "start": 8,
      "end": 9,
      "id": 1,
      "ws": true
    },
...
    {
      "text": ".",
      "start": 612,
      "end": 613,
      "id": 163,
      "ws": false
    }
  ],
  "_view_id": "ner_manual",
  "answer": "accept",
  "_timestamp": 1673378743,
  "spans": [
    {
      "start": 12,
      "end": 44,
      "token_start": 3,
      "token_end": 11,
      "label": "dosage",
      "text": "241 6 - 375mg split into 3 doses"
    },
    {
      "start": 192,
      "end": 224,
      "token_start": 56,
      "token_end": 64,
      "label": "dosage",
      "text": "241 4 - 375mg split into 3 doses"
    }
  ]
}

```

---

<div class="post-metadata">

**Author:** ![stefan.bartell](https://avatars.discourse-cdn.com/v4/letter/s/3d9bf3/32.png) [@stefan.bartell](https://support.prodi.gy/u/stefan.bartell)\
**Post date:** [January 11, 2023, 9:40pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/9 "2023-01-11T21:40:50Z")

</div>

Looks like it's working for you! I guess it is something small that is still causing me trouble. I'm getting `File ...line 40, in db_out for span in eg["spans"]: KeyError: 'spans'`

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [January 11, 2023, 9:47pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/10 "2023-01-11T21:47:57Z")

</div>

Ah, sorry! Completely forgot. You get this if you have a record that can't find a `spans`.

Change to this (I've also updated the code above):

```python
    for eg in examples:
        if eg.get('spans') is not None:
            for span in eg.get('spans'):
                span['text'] = eg['text'][span['start']:span['end']]

```

By using `eg.get('spans')` instead of `eg["spans"]` you won't get an error when it doesn't find a key.

Crossing fingers that this should work 🤞

---

<div class="post-metadata">

**Author:** ![stefan.bartell](https://avatars.discourse-cdn.com/v4/letter/s/3d9bf3/32.png) [@stefan.bartell](https://support.prodi.gy/u/stefan.bartell)\
**Post date:** [January 11, 2023, 9:52pm UTC](https://support.prodi.gy/t/empty-spans-and-spans-with-no-text-attribute/6224/11 "2023-01-11T21:52:03Z")

</div>

Looks like it's working! Thanks for all your help!
