# Classifying long-documents based on small spans of text

**URL:** <https://support.prodi.gy/t/classifying-long-documents-based-on-small-spans-of-text/3829>\
**Category:** Uncategorized\
**Tags:** usage, textcat, medical\
**Created:** [January 26, 2021, 7:55pm UTC](https://support.prodi.gy/t/classifying-long-documents-based-on-small-spans-of-text/3829 "2021-01-26T19:55:59Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [January 28, 2021, 1:02am UTC](https://support.prodi.gy/t/classifying-long-documents-based-on-small-spans-of-text/3829/2 "2021-01-28T01:02:18Z")

</div>

Hi, thanks so much for the kind words! Glad to hear that Prodigy has been useful so far 😊 This is also a very interesting and relevant use case, so definitely keep us updated on your progress!

> [@beperron](#):
>
> Now, I am wanting to classify the documents as present/absent for an opioid-problem. Does it make sense to first filter out all documents that do not contain any of the opioid-related terms?

That sounds like a good plan, yes, and it's definitely something I would try! If you have a reliable and reasonably accurate process to identify the opioid-related terms (terminology lists, named entities etc.), you can make the text classification task much more specific and potentially much more accurate because it doesn't _also_ have to learn whether a text is even relevant in the first place. You just want to make sure that you apply the same selection process during annotation and at runtime. So when you process all of your summaries later on, your workflow would be: check if the `doc` is relevant (contains related terms), then check the `doc.cats` for the predicted `HAS_PROBLEM` score (or something like that).

> [@beperron](#):
>
> For example, could I feed Prodigy the entire document or sentences, and Prodigy would highlight all the opioid-terms to facilitate document- or sentence-level annotation?

Yes, you could, for instance, use a simple custom recipe with the binary `classification` UI and stream in examples that contain a key `"spans"` describing the terms found in the example. You could also only send examples for annotation that contain terms, so you're focusing only on texts that are potentially relevant. See here for an example of the UI and JSON format: [Annotation interfaces · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/api-interfaces#classification)

Here's a simple example of how your stream logic could look: I've used spaCy's `PhraseMatcher` for matching the related terms. When you go through the texts, you can then check if a text is relevant (contains matches), add the matches as highlighted spans and send it out for annotation with the label, e.g. `HAS_PROBLEM`. For each example, you can then hit accept or reject.

```python
import spacy
from spacy.matcher import PhraseMatcher
from spacy.util import filter_spans

YOUR_TERMS = ["heroin", "methadone", "opioid"] # etc.
LABEL = "HAS_PROBLEM"

nlp = spacy.blank("en")
matcher = PhraseMatcher(nlp.vocab, attr="LOWER")
matcher.add("OPIOID", [nlp.make_doc(term) for term in YOUR_TERMS])

def make_stream(stream):
    # This expects a stream of examples like {"text": "..."}
    for doc in nlp.pipe((eg["text"] for eg in stream)):
        # Check if text is relevant, i.e. if it contains terms and only send out if relevant
        matches = matcher(doc)
        if matches: 
            matched_spans = [doc[start:end] for _, start, end in matches]
            # Just in case you have overlapping matches
            matched_spans = filter_spans(matched_spans)
            # Generate example and send out for annotation
            eg = {"text": doc.text, "spans": spans, "label": LABEL}
            yield eg

```

In a [custom recipe](https://prodi.gy/docs/custom-recipes), this could look like this:

```python
import prodigy

@prodigy.recipe("textcat.custom")
def textcat_custom(dataset, source):
     # Usage: prodigy textcat.custom dataset_name file.jsonl -F recipe.py
    stream = JSONL(source) # or however else you want to load the data
    stream = make_stream(stream) # function from above
    return {
        "dataset": dataset, # dataset to save annotations to
        "view_id": "classification", # UI to use
        "stream": stream, # data to stream in
    }

```

(You could make this a lot fancier if you feel like it – for example, define some more [recipe arguments](https://prodi.gy/docs/custom-recipes#recipe-args) so you can load in your terms from a file or pass in a label via the CLI.)

---

_[View the full topic](https://support.prodi.gy/t/classifying-long-documents-based-on-small-spans-of-text/3829)._
