# text classification

**URL:** <https://support.prodi.gy/t/text-classification/2063>\
**Category:** Uncategorized\
**Tags:** usage, textcat\
**Created:** [September 30, 2019, 5:44pm UTC](https://support.prodi.gy/t/text-classification/2063 "2019-09-30T17:44:10Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![robertto](https://avatars.discourse-cdn.com/v4/letter/r/f475e1/32.png) [@robertto](https://support.prodi.gy/u/robertto)\
**Post date:** [September 30, 2019, 5:44pm UTC](https://support.prodi.gy/t/text-classification/2063/1 "2019-09-30T17:44:10Z")

</div>

Hi guys,

after a successful NER model using a wonderful prodigy, I want to do content classification regarding observational data in my data set ( I have sentences I want to either annotate them as observational sents or edit my test labels)

I have provided jsonl data using this script

```python
texts = sents  
examples = []
options= [
        {"id": 0, "text": "negative"},
        {"id": 1, "text": "positive"}]
       
for text in texts:
    task = {"text": text,"options":options}
    examples.append(task)  
    
write_jsonl("dfObs_01.jsonl", examples)

```

I want to either annotate data by your interface or edit annotation, for the former one I used this command

```python
! python -m prodigy db-in dfObs dfObs_01.jsonl

```

then in order to classify to see the labels, I have done this

```python
!python -m prodigy textcat.teach df_obs en_core_web_sm --label Observation

```

but interface goes to loading and does not show the sentences, am I correct track?  
if I want to use my label (y) then modify it , how can I do it?  
I am familiar with NER and want to kind of do the similar with text classification (edit label, make model) I would appreciate your response

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [October 1, 2019, 11:35am UTC](https://support.prodi.gy/t/text-classification/2063/2 "2019-10-01T11:35:37Z")

</div>

Hi! I think the problem here is that you've imported your data to a Prodigy dataset, which holds the **collected annotations**. I think what you want to do instead is that your JSONL file and load it in as the source you're annotating in `textcat.teach`, e.g. as the third argument:

```bash
!python -m prodigy textcat.teach new_dataset en_core_web_sm dfObs_01.jsonl --label Observation

```

What do you want to do with your `options`? The `textcat.teach` recipe will only render the text with a given label, so if you want to have multiple-choice options, it sounds like the `textcat.manual` workflow is a better idea?

Also make sure you're using a new dataset to save the annotations to (otherwise, they'll be added to your previous data, which makes things messy). The reason you only saw "Loading..." btw is that Prodigy supports leaving out the `source` argument and piping data forward from a previous process. Because no source was set, it was basically waiting to read data from standard input.

---

<div class="post-metadata">

**Author:** ![robertto](https://avatars.discourse-cdn.com/v4/letter/r/f475e1/32.png) [@robertto](https://support.prodi.gy/u/robertto)\
**Post date:** [October 1, 2019, 1:01pm UTC](https://support.prodi.gy/t/text-classification/2063/3 "2019-10-01T13:01:20Z")

</div>

I tried to start to make workflow as following.

Work on my jsonl data (without label) by this:

```python
python -m prodigy textcat.manual dfObsV0003 en_core_web_sm dfObsV03.jsonl --label Observation
Using 1 labels: Observation

```

but I faced with this error:

```python

✨ ERROR: Invalid task format for view ID 'classification'
'label' is a required property

{'text': 'Chapter 1', '_input_hash': 1891558552, '_task_hash': -2011871074, '_session_id': 'dfObsV0003-default', '_view_id': 'classification'}

```

as you You mentioned here

> [@Only 25 lines loading from my .jsonl stream](https://support.prodi.gy/t/only-25-lines-loading-from-my-jsonl-stream/1838/2):
>
> Hi! One thing to keep in mind when using the active learning-powered recipes like textcat.teach is that they don’t necessarily show you all examples you’re loading in. The main concept behind the active learning approach is to show you the most relevant examples for annotation using the model’s predictions. Under the hood, Prodigy uses an exponential moving average of the scores to decide whether to send an example out for annotation or not. So based on the annotation decisions you make, the m…

that is bug. I add the edited version of script (since it has indent error)

```python
def add_label_to_stream(stream, label):
        for eg in stream:
            eg["label"] = label[0]
            yield eg
        if has_options:
            stream = add_label_options(stream, label)
        else:
            stream = add_label_to_stream(stream, label)

```

to end of texcat. it does not work I also add that to end of  
recipe manual...again it does not work...!  
Am I doing any mistake? could be related to my jsonl data? can you give me a step by step way to manually annotate my data and then make a model based on my annotation?

---

<div class="post-metadata">

**Author:** ![robertto](https://avatars.discourse-cdn.com/v4/letter/r/f475e1/32.png) [@robertto](https://support.prodi.gy/u/robertto)\
**Post date:** [October 2, 2019, 8:01am UTC](https://support.prodi.gy/t/text-classification/2063/4 "2019-10-02T08:01:41Z")

</div>

Hi @ines  
Morning! good news! I managed to run  
textcat.manual on a raw text without label in order to gather labels as follows

```python
python -m prodigy textcat.manual dfobsv02 en_core_web_sm dfObsV02.jsonl --label Observational,Nonobservational --exclusive

```

now annotator can manually annotate observational sentence from nonobservational sentences

my question is if I have predefined labels (y values, basically 0 and 1) how can I run textcat.manual to edit those (similar to ner.manual)  
I think I should change the type of my data since now it is so:

```python
{"text":"my text."}
{"text":"my test 02."}
{"text":"my test 03."}

```

my second question is , can I see my ner labels when I classify the sentences in textcat.manual how to use my NER label .. to kind of improve classification?

and my last question maybe it is not related to this threat. how can I add-relation to entities?  
I have read this

> [@Relation Extraction annotation tests](https://support.prodi.gy/t/relation-extraction-annotation-tests/639):
>
> Hi @ines, In this other [thread](https://support.prodi.gy/t/annotation-for-argument-mining/601/14), you told me the following: When you’re done with that, you could export the data and run one experiment where you link up highlighted spans close to each other and collect binary feedback on whether they are connected. You could either create data in the dep format (with a head and child), or use the choice interface with the options “For”, “Against”, “Support” and “Attack”. So I was testing the two approaches, but I’m having a hard time making them work. I’…

still, dot now how can I use my annotated data (ner) by prodigy to make training set for relation extraction

I really like the prodigy , I have learned a lot here.  
Best

---

<div class="post-metadata">

**Author:** ![robertto](https://avatars.discourse-cdn.com/v4/letter/r/f475e1/32.png) [@robertto](https://support.prodi.gy/u/robertto)\
**Post date:** [October 7, 2019, 10:36am UTC](https://support.prodi.gy/t/text-classification/2063/5 "2019-10-07T10:36:06Z")

</div>

hi,

I would be very thankful if you can answer my question in last comment?

many many thanks

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [October 7, 2019, 12:07pm UTC](https://support.prodi.gy/t/text-classification/2063/6 "2019-10-07T12:07:54Z")

</div>

> [@robertto](#):
>
> my question is if I have predefined labels (y values, basically 0 and 1) how can I run textcat.manual to edit those (similar to ner.manual)

I'm a bit confused by what you're trying to do. Do you want to assign both top-level categories and highlight spans in the text? If so, you probably want to do this in two different steps.

> [@robertto](#):
>
> can I see my ner labels when I classify the sentences in textcat.manual how to use my NER label .. to kind of improve classification?

Do you mean, use the predicted entity spans as features in the text classifier? Not by default if you're using spaCy text classification implementation. You'd probably have to build something custom. Although, the entities being present in the text will likely still have an impact – if texts containing entity X are typically about Y, the text classifier can pick up on that based on the words occuring in the text.

> [@robertto](#):
>
> still, dot now how can I use my annotated data (ner) by prodigy to make training set for relation extraction

This depends on what data you need to train your relation extraction model. One approach could be to stream in pairs of entities that are close in the text and then annotate their relations using the `choice` interface, as described in the thread you linked.

---

<div class="post-metadata">

**Author:** ![robertto](https://avatars.discourse-cdn.com/v4/letter/r/f475e1/32.png) [@robertto](https://support.prodi.gy/u/robertto)\
**Post date:** [October 7, 2019, 3:47pm UTC](https://support.prodi.gy/t/text-classification/2063/7 "2019-10-07T15:47:52Z")

</div>

"  
I'm a bit confused by what you're trying to do. Do you want to assign both top-level categories and highlight spans in the text? If so, you probably want to do this in two different steps.  
"

I want to use the classification interface (texcat) and also see the result of my NER in the monitor.is it possible?

"  
This depends on what data you need to train your relation extraction model. One approach could be to stream in pairs of entities that are close in the text and then annotate their relations using the `choice` interface, as described in the thread you linked.  
"

how can I start this choice interface ?

* * *

another different idea:

I want to create large realtion (inclusing causally)-labelled datasets for supervised machine learning NP1-VERB-NP2

I have created the verbs by reverb instruction

 ![re_daivd_Capture](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/0/0eef5dcecd10ba3546b0b5762f62c68c1acedd33.png)

then I want to provide a dataset like this:

> **[da551ef6b1c909ca5b37ba94be4cae02e9ac.pdf](https://pdfs.semanticscholar.org/8ea8/da551ef6b1c909ca5b37ba94be4cae02e9ac.pdf)**
>
> 133.17 KB

FUNCTION relation between airplane and transportation in “the airplane  
is used for transportation”

a PART-WHOLE relation in “the car has an engine”.

ACQUISITION between named entities in “Yahoo has made a definitive agreement to  
acquire Flickr”.

here, There is a bit missing which is to find the closest noun-phrases to right and left of the pattern which gave me the verb, then I want to kind annotate each pair in the sentences to different relation as well as  
"causal relation ", "FUNCTION relation"or....

for this aim, I need first find noun-phrases to right and left of the pattern (verbs)

then I have kind of

```python
"text" , "e1","e2" ,"relation",

```

I want to assign this relation in prodigy,then train a model on that 🙂

my question how can I start in prodigy?

do you have any hints how can I find closet name in right and left to verb ?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [October 7, 2019, 6:01pm UTC](https://support.prodi.gy/t/text-classification/2063/8 "2019-10-07T18:01:07Z")

</div>

> [@robertto](#):
>
> I want to use the classification interface (texcat) and also see the result of my NER in the monitor.is it possible?

Have you tried streaming in data that contains the `"spans"`? I think `textcat.teach` might reset the existing spans because it also uses spans to pre-highlight pattern matches. But if you're not using a model in the loop and are just labelling with multiple-choice options, you can just leave the pre-anotated spans in the data, add the options and you should see them highlighted in the interface.

> [@robertto](#):
>
> how can I start this choice interface ?

Check out the documentation on custom recipes and interfaces, for example, starting here: [Custom Recipes · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/workflow-custom-recipes#example-choice) You might also want to check out the `PRODIGY_README.html`, which includes the detailed documentation of the components.

> [@robertto](#):
>
> for this aim, I need first find noun-phrases to right and left of the pattern (verbs)

You could start by labelling all noun phrases you're interested in, either by hand or using spaCy to pre-select them for you (e.g. via the `Doc.noun_chunks` or just by extracting noin tokens). This will give you the spans of tokens and their position in the text. You could then use spaCy to extract the verbs attached to the nouns – e.g. using [the dependency parse](https://spacy.io/usage/linguistic-features#dependency-parse). Next, you can stream this information into Prodigy and use the choice interface to select the relation – e.g. using the `choice` interface. You might have to experiment a bit to see what works best.
