# Prodigy error when reviewing audio annotation coupled with videos

**URL:** <https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627>\
**Category:** Uncategorized\
**Tags:** usage, audio, video\
**Created:** [November 12, 2020, 4:57pm UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627 "2020-11-12T16:57:32Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Gnai](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/gnai/32/1637_2.png) [@Gnai](https://support.prodi.gy/u/Gnai)\
**Post date:** [November 12, 2020, 4:57pm UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/1 "2020-11-12T16:57:32Z")

</div>

Hello,  
We are trying to review some audio annotations done by labelers. After encoding our `jsonl` file with the right data to `base 64`, we ended up with a 5gb encoded `jsonl` file supposedly for ~80 videos.

Running this locally with `cat ~/audio_b64.jsonl | prodigy audio.manual rev_audio - --loader jsonl`  
`--label person`  
couldn't load our file for revieweing the annotations, and prompted with this in our terminal:

```python
⚠ Warning: filtered 99% of entries because they were duplicates. Only 1
items were shown out of 77. You may want to deduplicate your dataset ahead of
time to get a better understanding of your dataset size.

```

Any idea what might be the cause of that? What can be an alternative?  
Thank you.  
George.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [November 20, 2020, 2:51am UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/2 "2020-11-20T02:51:54Z")

</div>

Hi and sorry for only getting to this now, I must have missed the thread!

How does your JSONL data look under the hood? It sounds like the hashes Prodigy auto-generated for it aren't taking the actual video data into account, so all records end up with the same hashes. When generating the JSONL, you add an `_input_hash` and `_task_hash` value (e.g. based on the file name) to help Prodigy distinguish between examples that are identical, different questions about the same data, and entirely different inputs/questions.

Btw, if your data is this large, you might also want to consider a different loading strategy so you don't end up with these huge JSONL files and base64 strings that are sent back and forth. One option could be to use an S3 bucket (or similar) and only have your JSONL contain the URLs.

---

<div class="post-metadata">

**Author:** ![Gnai](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/gnai/32/1637_2.png) [@Gnai](https://support.prodi.gy/u/Gnai)\
**Post date:** [November 24, 2020, 4:51pm UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/3 "2020-11-24T16:51:46Z")

</div>

Hello Ines,  
Yes turns out under the hood my `jsonl` file have actually the same `_input_hash` only the `_task_hash` are different.  
I tried dropping the `_input_hash` and changing the number by one or two digits so it differs a bit, in both cases it was still raising the same error. Any ideas on what to do next?

I switched to reading from the url directly, thank you for that, I also thought it was maybe a problem of encoding to base 64.

Thank you.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [November 25, 2020, 12:39am UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/4 "2020-11-25T00:39:04Z")

</div>

> [@Gnai](#):
>
> I tried dropping the `_input_hash` and changing the number by one or two digits so it differs a bit, in both cases it was still raising the same error. Any ideas on what to do next?

Okay, so you mean, you still saw the warning about 99% of entries being filtered? I just double-checked and the `audio` recipe actually re-hashes the stream by default, and for some reason, it doesn't seem to take the `"video"` key into account.

In the meantime, can you try assigning a unique `"text"` value to each entry? This should do the trick and convince Prodigy that the records are all different.

---

<div class="post-metadata">

**Author:** ![Gnai](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/gnai/32/1637_2.png) [@Gnai](https://support.prodi.gy/u/Gnai)\
**Post date:** [November 25, 2020, 11:52am UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/5 "2020-11-25T11:52:23Z")

</div>

Hello Ines,  
For the `text` and `_input_hash` values I assigned different values respectively for each entry. I am trying with a small `jsonl` file now to see.  
Should an entry in my file be looking something like this? where should I put the url link? -- Keep in mind that I am trying to run this locally first and then will try it on an EC2 instance where we actually have the data as well.

```python
{"text":"e","meta":{"path":"/Users/macuser/Documents/annotations/av_ann/PC_25112020/input_1_17.mp4"},"file":"","_input_hash":614455852,"_task_hash":1401296448,"_session_id":"laudio_prodigy_data","_view_id":"blocks","audio_spans":[{"start":-0.0100006104,"end":66.1899993896,"label":"LEAD_SINGER_M","id":"b3837d9b-baec-43b1-8287-3f160ad52fd1","color":"rgba(255,215,0,0.2)"}],"answer":"accept","ts":"2020-11-24-07","audio_clarity":4.0,"loudness":4.0,"noisiness":4.0}

```

Thanks a lot.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [November 26, 2020, 12:35am UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/6 "2020-11-26T00:35:45Z")

</div>

That looks good, but you still need a key `"video"` or `"audio"` in your entry that includes the base64-encoded data or a URL to the file (can't be a local path, though, because your browser will most likely block this for security reasons).

---

<div class="post-metadata">

**Author:** ![Gnai](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/gnai/32/1637_2.png) [@Gnai](https://support.prodi.gy/u/Gnai)\
**Post date:** [December 4, 2020, 3:47pm UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/7 "2020-12-04T15:47:37Z")

</div>

Hello Ines,

If I understood correctly if I connect that to my ec2 data, that would be blocked by my browser?  
I was trying to fetch the video files from my s3 by using `boto3` and `botocore` but no luck, can you point out an example somewhere?  
From what I have understood, that would change the `loaders` in our command line as well, right?  
So eventually our command line should be looking something like that:  
`python3 -m prodigy audio_custom.manual testdb --label singer -F audio_recipe_s3.py -`

Thank you for the support.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [December 5, 2020, 4:51am UTC](https://support.prodi.gy/t/prodigy-error-when-reviewing-audio-annotation-coupled-with-videos/3627/8 "2020-12-05T04:51:33Z")

</div>

> [@Gnai](#):
>
> If I understood correctly if I connect that to my ec2 data, that would be blocked by my browser?

If the image path is a http(s) URL, that's fine – but your browser would likely block paths like `/users/you/some_file.jpg` (referring to local paths on your filesystem) and you typically don't want to disable this behaviour either. This thread has some more background on this:

> <https://stackoverflow.com/questions/39007243/cannot-open-local-file-chrome-not-allowed-to-load-local-resource>

> [@Gnai](#):
>
> From what I have understood, that would change the `loaders` in our command line as well, right?

If you're using one of the built-in recipes, you can set the `--loader` argument to define how the file should be loaded. So if you want to load your videos from a JSONL file containing URLs, you'd set `--loader jsonl`.

In a custom recipe, you can define the command line arguments however you like. See here for details and examples: [Custom Recipes · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/custom-recipes#recipe-args)
