# Textcat not excluding dataset.

**URL:** <https://support.prodi.gy/t/textcat-not-excluding-dataset/3101>\
**Category:** Uncategorized\
**Tags:** textcat, streams\
**Created:** [June 28, 2020, 7:09pm UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101 "2020-06-28T19:09:02Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tim](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tim/32/1500_2.png) [@Tim](https://support.prodi.gy/u/Tim)\
**Post date:** [June 28, 2020, 7:09pm UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/1 "2020-06-28T19:09:02Z")

</div>

I was working on manual textcat (using textcat.manual) yesterday and had to stop, so I saved my work assuming I could exclude a existing dataset, which according to the documentation is possible. However when I run the following command, it does not exclude the data in my original dataset "cats\_m-1" and still asks me the questions I already answered in "cats\_m-1"

```
prodigy textcat.manual cats_m-f2 data\all_data.jsonl --label <Long list of labels here> -E -e cats_m-1

```

I could just write a script that automatically removes the data in my original dataset from my jsonl file, but this seems unnecessary when there is a special function to exclude a existing dataset.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 29, 2020, 8:22am UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/2 "2020-06-29T08:22:53Z")

</div>

Hi! That workflow sounds reasonable and your workflow looks correct. By default, the `textcat.manual` recipe with options will exclude based on the input hash (representing the text), so if a question with the same input hash already exists, it should skip it.

If you look at the `_input_hash` values in the data, are they identical? And what does it say about excluding in the logs when you set `PRODIGY_LOGGING=basic`?

---

<div class="post-metadata">

**Author:** ![Tim](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tim/32/1500_2.png) [@Tim](https://support.prodi.gy/u/Tim)\
**Post date:** [June 29, 2020, 11:34am UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/3 "2020-06-29T11:34:55Z")

</div>

Hey.

> If you look at the `_input_hash` values in the data, are they identical?

They indeed are.

> And what does it say about excluding in the logs when you set `PRODIGY_LOGGING=basic` ?

I've not been able to figure out how to enable this. I looked in the documentation, and it simply tells me to set a environment variable. I am unable to find where this is located.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 29, 2020, 12:58pm UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/4 "2020-06-29T12:58:47Z")

</div>

> [@Tim](#):
>
> I've not been able to figure out how to enable this. I looked in the documentation, and it simply tells me to set a environment variable. I am unable to find where this is located.

Ah, so this is just a regular environment variable that you can set however you'd normally set an environment variable on your platform (differs by platform, dev environment etc.).

If you're on a Mac or on Linux, you can do the following. (Also see here for an example: [Installation & Setup · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/install#debugging-logging))

```bash
PRODIGY_LOGGING=basic prodigy textcat.manual cats_m-f2 data\all_data.jsonl --label <Long list of labels here> -E -e cats_m-1

```

---

<div class="post-metadata">

**Author:** ![Tim](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tim/32/1500_2.png) [@Tim](https://support.prodi.gy/u/Tim)\
**Post date:** [June 29, 2020, 4:49pm UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/5 "2020-06-29T16:49:25Z")

</div>

> [@Tim](#):
>
> ```python
> prodigy textcat.manual cats_m-f2 data\all_data.jsonl --label <Long list of labels here> -E -e cats_m-1
> 
> ```

Got it, thank you.  
Log below:

18:47:04: RECIPE: Starting recipe textcat.manual  
18:47:04: RECIPE: Annotating with 9 labels  
18:47:04: LOADER: Using file extension 'jsonl' to find loader  
18:47:04: LOADER: Loading stream from jsonl  
18:47:04: LOADER: Rehashing stream  
18:47:04: VALIDATE: Validating components returned by recipe  
18:47:04: CONTROLLER: Initialising from recipe  
18:47:04: VALIDATE: Creating validator for view ID 'choice'  
18:47:04: VALIDATE: Validating Prodigy and recipe config  
18:47:04: DB: Initializing database SQLite  
18:47:04: DB: Connecting to database SQLite  
18:47:04: DB: Creating dataset '2020-06-29\_18-47-04'  
18:47:04: CONTROLLER: Initialising from recipe  
18:47:04: CONTROLLER: Validating the first batch for session: None  
18:47:04: PREPROCESS: Add multiple choice options for 9 labels  
18:47:04: FILTER: Filtering duplicates from stream  
18:47:04: FILTER: Filtering out empty examples for key 'text'  
18:47:04: CORS: initialized with wildcard "\*" CORS origins

Personally can't see anything related to excluding a dataset.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 29, 2020, 4:56pm UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/6 "2020-06-29T16:56:26Z")

</div>

Cool, glad it worked! This looks like it's just the log on startup – the filtering happens when the stream is loaded and processed, so you might just have to open the app in the browser so it starts queuing up some examples.

---

<div class="post-metadata">

**Author:** ![Tim](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/tim/32/1500_2.png) [@Tim](https://support.prodi.gy/u/Tim)\
**Post date:** [June 29, 2020, 5:00pm UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/7 "2020-06-29T17:00:29Z")

</div>

Ok, I've done that. It's now showing the following.

INFO: ::1:52278 - "GET / HTTP/1.1" 200 OK  
INFO: ::1:52278 - "GET /bundle.js HTTP/1.1" 200 OK  
18:56:46: GET: /project  
INFO: ::1:52278 - "GET /project HTTP/1.1" 200 OK  
18:56:46: POST: /get\_session\_questions  
18:56:46: FEED: Finding next batch of questions in stream  
18:56:46: RESPONSE: /get\_session\_questions (10 examples)  
INFO: ::1:52278 - "POST /get\_session\_questions HTTP/1.1" 200 OK  
18:57:17: POST: /get\_session\_questions  
18:57:17: FEED: Finding next batch of questions in stream  
18:57:17: FEED: skipped: 1884547963 (text)  
18:57:17: FEED: skipped: -980977358 (text)  
18:57:17: FEED: skipped: 1450570456 (text)  
18:57:17: FEED: skipped: 1842367882 (text)  
18:57:17: FEED: skipped: 981677283 (text)  
18:57:17: FEED: skipped: -2003029328 (text)  
18:57:17: FEED: skipped: -1195566099 (text)  
18:57:17: FEED: skipped: 966435862 (text)  
18:57:17: RESPONSE: /get\_session\_questions (10 examples)  
INFO: ::1:52279 - "POST /get\_session\_questions HTTP/1.1" 200 OK

For some reason, I think it's working right now. Which is weird because my command is exactly the same, the only difference is that I enabled logging.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 30, 2020, 11:43am UTC](https://support.prodi.gy/t/textcat-not-excluding-dataset/3101/8 "2020-06-30T11:43:07Z")

</div>

> [@Tim](#):
>
> For some reason, I think it's working right now. Which is weird because my command is exactly the same, the only difference is that I enabled logging.

Glad to hear it's working! Still a little strange, though, because the logging just toggles Python's logging and I don't see how there could be any interaction there 🤔 Anyway, if it happens again let me know. The only possible theory I can think of is that for _some_ reason, the `-e cats_m-1` at the end of your command may have not been interpreted correctly... but then again, this would have likely caused an error.
