# Simple way of getting tagged/marked words and phrases after span categorization task

**URL:** <https://support.prodi.gy/t/simple-way-of-getting-tagged-marked-words-and-phrases-after-span-categorization-task/6230>\
**Category:** Uncategorized\
**Tags:** solved\
**Created:** [January 12, 2023, 2:23am UTC](https://support.prodi.gy/t/simple-way-of-getting-tagged-marked-words-and-phrases-after-span-categorization-task/6230 "2023-01-12T02:23:42Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![joebuckle](https://avatars.discourse-cdn.com/v4/letter/j/ee7513/32.png) [@joebuckle](https://support.prodi.gy/u/joebuckle)\
**Post date:** [January 12, 2023, 2:23am UTC](https://support.prodi.gy/t/simple-way-of-getting-tagged-marked-words-and-phrases-after-span-categorization-task/6230/1 "2023-01-12T02:23:42Z")

</div>

Good day,

Is there a simple way to get all the words and phrases that we have tagged in a span categorization task? We are using this task to be able to get common keywords/phrases in the text.

Regards,  
Joe

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [January 12, 2023, 3:15pm UTC](https://support.prodi.gy/t/simple-way-of-getting-tagged-marked-words-and-phrases-after-span-categorization-task/6230/2 "2023-01-12T15:15:36Z")

</div>

Hi Joe,

could you clarify what you mean with "get all the words and phrases that we have tagged"? You could use the [`db-out`](https://prodi.gy/docs/recipes#db-out) command but I'm not 100% if that's what you mean.

---

<div class="post-metadata">

**Author:** ![joebuckle](https://avatars.discourse-cdn.com/v4/letter/j/ee7513/32.png) [@joebuckle](https://support.prodi.gy/u/joebuckle)\
**Post date:** [January 12, 2023, 10:43pm UTC](https://support.prodi.gy/t/simple-way-of-getting-tagged-marked-words-and-phrases-after-span-categorization-task/6230/3 "2023-01-12T22:43:32Z")

</div>

Sorry, I mean to be able to get the actual words and phrases within a custom recipe (python code). Because looking at the example on 'accept', it contains only the start and end numbers of the tokens.

Thanks.

---

<div class="post-metadata">

**Author:** ![koaning](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/koaning/32/230_2.png) [@koaning](https://support.prodi.gy/u/koaning)\
**Post date:** [January 13, 2023, 10:02am UTC](https://support.prodi.gy/t/simple-way-of-getting-tagged-marked-words-and-phrases-after-span-categorization-task/6230/4 "2023-01-13T10:02:49Z")

</div>

You could fetch those via the `update` callback. Here's a custom recipe with a demo.

```python
from pathlib import Path
from typing import Any, Dict

import prodigy
import spacy
from prodigy import get_stream
from prodigy.components.preprocess import add_tokens

def update(answers):
    for answer in answers:
        for span in answer['spans']:
            start = span['start']
            end = span['end']
            print(f"I found this span: {answer['text'][start:end]}")

@prodigy.recipe(
    "spancat.special",
    dataset=("Dataset to save annotations to", "positional", None, str),
    lang=("language for the tokeniser", "positional", None, str),
    source=("Data to annotate", "positional", None, str),
    labels=("comma separated sequence of labels", "option", "l", str),
)
def special(dataset: str, lang: str, source: Path, labels: str) -> Dict[str, Any]:
    nlp = spacy.blank(lang)
    labels = labels.split(",")
    stream = get_stream(source, rehash=True, input_key="text", dedup=True)
    stream = add_tokens(nlp, stream, skip=True)

    return {
        "dataset": dataset,
        "stream": stream,
        "view_id": "spans_manual",
        "update": update,
        "config": {
            "lang": lang,
            "labels": labels,
            "batch_size": 1
        },
    }

```

This recipe is in a folder with an examples.jsonl file that contains:

```python
{"text": "hi my name is Vincent D. Warmerdam"}
{"text": "hi my name is Johnny Bravo"}

```

When I run this command:

```python
python -m prodigy spancat.special issue-6230 en examples.jsonl --labels name -F recipe.py

```

I get this annotation interface.

 ![CleanShot 2023-01-13 at 11.01.19](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/3/3b6c0dce860f81041a15302243bdb5fff0b399a9.png)

It's very much just the spancat interface, but try to annotate a single example and hit "save". When you do, you should see a print appear with the selected span. Does this suffice?

---

<div class="post-metadata">

**Author:** ![joebuckle](https://avatars.discourse-cdn.com/v4/letter/j/ee7513/32.png) [@joebuckle](https://support.prodi.gy/u/joebuckle)\
**Post date:** [January 14, 2023, 3:19am UTC](https://support.prodi.gy/t/simple-way-of-getting-tagged-marked-words-and-phrases-after-span-categorization-task/6230/5 "2023-01-14T03:19:09Z")

</div>

Yes, thank you very much!
