# Labelling a set of images (classification)

**URL:** <https://support.prodi.gy/t/labelling-a-set-of-images-classification/4608>\
**Category:** Uncategorized\
**Tags:** usage, image\
**Created:** [August 23, 2021, 7:40pm UTC](https://support.prodi.gy/t/labelling-a-set-of-images-classification/4608 "2021-08-23T19:40:52Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [August 24, 2021, 2:39am UTC](https://support.prodi.gy/t/labelling-a-set-of-images-classification/4608/2 "2021-08-24T02:39:12Z")

</div>

Hi! It sounds like you're definitely on the right track 🙂

Instead of using the `mark` recipe, which really just streams in what you give it, you might actually find it easier to just implement a custom recipe for this, since it'll make it more obvious what's going on and lets you add your own custom logic (e.g. for shuffling, removing base64 and maybe other stuff).

This example recipe actually goes in a very similar directon: [Computer Vision · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/computer-vision#classification-multi) – only that in your case, you'd add a single `"label"` to the examples instead of `"options"`, and use the `classification` interface instead of `choice`.

> [@strickvl](#):
>
> Q: Is there a way to have images loaded in randomly? (I reckon one way of doing this would be to rename all the files with random alphanumeric strings, though then I'd lose the original file names. Is there another way?)

Sure, that's definitely reasonable. When you load your images from a directory using the `Images` loader, what you get back is a regular Python generator that yields dictionaries:

```python
stream = Images(source)

```

The most straightforward solution would be to just call `list()` and `random.shuffle()` on it to shuffle it – however, this will consume the whole generator upfront. Another option would be to go through your stream and use some heuristic to (randomly) decide whether or not to send out a given example for annotation. For instance:

```python
def get_random_stream():
    stream = Images(source)
    for eg in stream:
        if random.random() > 0.7: # or whatever
            yield eg

```

> [@strickvl](#):
>
> Q: Is there a way to have Prodigy not ask me to relabel images that I've already labelled?

Feed overlap isn't what you want here, because that just controls whether multiple annotators in different sessions are asked about the same example or not.

By default, Prodigy will generate two hashes for each example: one representing the input (e.g. the image) and one representing the question about the image (e.g. image + label). If an example with the same task hash is already present in the current dataset, you shouldn't be asked about it again. So you'd see _different questions_ about the same image, but not the same questions about the same image. Alternatively, you can also set `"exclude_by": "input"` in the `"config"` returned by your recipe to exclude based on the input hash. In that case, you would only see a given image once.

If your images don't change between runs and you're saving your annotations to the same dataset, you should only be asked about images you haven't annotated yet.

> [@strickvl](#):
>
> Q: Is there a way to set `--remove-base64` when using the `mark` recipe?

You can do this in your custom recipe by adding a `before_db` callback, that can modify examples in place before they're added to the database. Here's an example of the same code the built-in image recipes use to implement `--remove-base64`: [Custom Recipes · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/custom-recipes#before_db)

This will replace the base64 string with the path, so you just need to make sure the files don't change. Definitely be careful here, though, because you don't want to accidentally destroy any data.

If you don't want to convert the images to base64, you can also use the `ImageServer` loader instead, which provides the image URLs via a local web server: [Loaders and Input Data · Prodigy · An annotation tool for AI, Machine Learning & NLP](https://prodi.gy/docs/api-loaders#loaders-file) Another alternative is to just put your images in an S3 bucket or similar and use the URLs instead.

---

_[View the full topic](https://support.prodi.gy/t/labelling-a-set-of-images-classification/4608)._
