# Use Prodigy purely as an annotating tool?

**URL:** <https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981>\
**Category:** Uncategorized\
**Tags:** usage, spacy, solved\
**Created:** [November 19, 2018, 4:18pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981 "2018-11-19T16:18:48Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![wsdee1](https://avatars.discourse-cdn.com/v4/letter/w/f07891/32.png) [@wsdee1](https://support.prodi.gy/u/wsdee1)\
**Post date:** [November 19, 2018, 4:18pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/1 "2018-11-19T16:18:48Z")

</div>

Hi there,

I am new to both prodigy and spacy, and I just wanted to use prodigy as an annotation tool. Say for example: I have a set of data which contains 10 entries, where 2 of them are false, and I need to filter out these 2 false entries. Finally, I wish to export the filtered result in whatever formats for other usage.

All the best.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [November 19, 2018, 5:35pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/2 "2018-11-19T17:35:50Z")

</div>

Hi! And yes, absolutely! Prodigy mostly orchestrates the flow of incoming data → annotation UI → callbacks → collected annotations. Each of these workflows is expressed via a “recipe”, a Python script that returns a dictionary of components. While the tool ships with various advanced recipes that update a model in the loop etc., you can also just stream in pretty much any data, visualize it, label it and get the result back.

For example, let’s say your input is a `.jsonl` file that looks like this:

```json
{"text": "Hello world"}
{"text": "Another text"}
{"text": "And another one"}

```

You could then use the `mark` recipe ([details](https://prodi.gy/docs/recipes#mark)), which takes whatever data comes in, presents it for annotation and saves the results in the dataset. For example:

```bash
prodigy mark your_dataset_name /path/to/data.jsonl --view-id text

```

The `view-id` is the name of the [annotation interface](https://prodi.gy/docs/web-app#interfaces) to use. Depending on the interface, you can add more properties to your data – e.g. the named entity spans, a top-level label, an image, HTML etc. You can find more details and examples of the formats in your `PRODIGY_README.html`.

Executing the above command will start Prodigy and serve up the annotation app. You can then open it in your browser and start labelling. Your answers will be sent back to the server and saved in the dataset. After annotation, you can export the collected data from the dataset:

```bash
prodigy db-out your_dataset_name > some_file.jsonl

```

The result could look like this:

```json
{"text": "Hello world", "answer": "accept"}
{"text": "Another text", "answer": "reject"}
{"text": "And another one", "answer": "ignore"}

```

Prodigy uses newline-delimited JSON as its standard output format, since it’s easy to work with and easy to read in and manipulate in any language or library. So if you need a different format, it shouldn’t be difficult to convert the data.

Btw, if you’re interested in writing your own recipe scripts that do more custom stuff (load in data from a different format, perform certain actions when you receive annotations, render something custom), a good place to start is the `prodigy-recipes` repo: [https://github.com/explosion/prodigy-recipes](https://github.com/explosion/prodigy-recipes) For example, here’s the code for the `mark` recipe I mentioned above:

> <https://github.com/explosion/prodigy-recipes/blob/master/other/mark.py>

---

<div class="post-metadata">

**Author:** ![wsdee1](https://avatars.discourse-cdn.com/v4/letter/w/f07891/32.png) [@wsdee1](https://support.prodi.gy/u/wsdee1)\
**Post date:** [November 22, 2018, 12:51pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/3 "2018-11-22T12:51:49Z")

</div>

Hi Ines,

So many thanks for your assistance, which is indeed, really helpful and quick!

All the best

---

<div class="post-metadata">

**Author:** ![alonisser](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/alonisser/32/552_2.png) [@alonisser](https://support.prodi.gy/u/alonisser)\
**Post date:** [December 4, 2018, 10:39pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/4 "2018-12-04T22:39:49Z")

</div>

We’ve followed this thread and run (where news\_headline\_mark is the dataset name we would like and news\_headlines\_options.jsonl an example of the news\_headline jsonl but with an options array)  
prodigy mark news\_headline\_mark news\_headlines\_options.jsonl --view-id choice  
we’ve been able to annotate with the UI, But we found out that results are saved into db only when we press the “save” button, This does not seems to be the behavior while using ner.teach instead of mark. How can we get “auto save” when the annotator is confirming a choice?

Thanks!

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [December 5, 2018, 1:24am UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/5 "2018-12-05T01:24:04Z")

</div>

> [@alonisser](#):
>
> This does not seems to be the behavior while using ner.teach instead of mark. How can we get “auto save” when the annotator is confirming a choice?

Annotations should be saved the same way across all recipes and interfaces. Prodigy sends collected annotations back to the server in batches, so as soon as a batch is full, the answers will be saved and sent back to the server automatically. The most recent answers are kept on the client to allow hitting "undo".

If you want the answers to be sent back sooner, you can change the `"batch_size"` setting in your `prodigy.json` or recipe config. The default batch size is `10`.

---

<div class="post-metadata">

**Author:** ![alonisser](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/alonisser/32/552_2.png) [@alonisser](https://support.prodi.gy/u/alonisser)\
**Post date:** [December 5, 2018, 7:23am UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/6 "2018-12-05T07:23:37Z")

</div>

Thanks. I’ll try that

---

<div class="post-metadata">

**Author:** ![alonisser](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/alonisser/32/552_2.png) [@alonisser](https://support.prodi.gy/u/alonisser)\
**Post date:** [December 9, 2018, 11:47am UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/7 "2018-12-09T11:47:23Z")

</div>

Thanks @ines I’ve tried reducing batch\_size to 5 in prodigy.json  
I fail to make it autosave. while using mark, if I don’t explicitly push “save” button , no matter if I annotated more then 5 items, it does not save to db  
Any solution to that?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [December 9, 2018, 12:09pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/8 "2018-12-09T12:09:24Z")

</div>

> [@alonisser](#):
>
> I fail to make it autosave. while using mark, if I don’t explicitly push “save” button , no matter if I annotated more then 5 items, it does not save to db

Hmm, that's strange! Initially, you do have to annotate two batches for the first one to get autosaved. The most recent annotations will always stay in the app so you can undo them easily – so before you close the tab, you'll always have to save manually to submit everything that's left.

One thing you can do to check for intermediate autosaving is open the developer tools and look at the console or network tab for requests made by the app. After the initial answers, the app should make a POST request to `/give_answers` (sending back one batch of answers). It will also periodically make requests to `/get_questions` to request new tasks from the server.

---

<div class="post-metadata">

**Author:** ![alonisser](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/alonisser/32/552_2.png) [@alonisser](https://support.prodi.gy/u/alonisser)\
**Post date:** [December 9, 2018, 12:42pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/9 "2018-12-09T12:42:58Z")

</div>

Oh, sorry, it’s was the initial “two batches” thing that got me. After the second batch, the first one was saved

---

<div class="post-metadata">

**Author:** ![Enix26](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/enix26/32/562_2.png) [@Enix26](https://support.prodi.gy/u/Enix26)\
**Post date:** [December 12, 2018, 1:45pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/10 "2018-12-12T13:45:10Z")

</div>

Hello Ines, Is it possible to customize the UI? Over what language is built?

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [December 12, 2018, 2:04pm UTC](https://support.prodi.gy/t/use-prodigy-purely-as-an-annotating-tool/981/11 "2018-12-12T14:04:13Z")

</div>

> [@Enix26](#):
>
> Hello Ines, Is it possible to customize the UI? Over what language is built?

The Prodigy library that configures the annotation workflows is written in Python, the app is build in JavaScript (React), shipped with the core library as a compiled bundle.

See this page for theming options and details on custom HTML annotation views:

> **[Web Application · Prodigy · An annotation tool for AI, Machine Learning &...](https://prodi.gy/docs/api-web-app/)**
>
> A downloadable annotation tool for NLP and computer vision tasks such as named entity recognition, text classification, object detection, image segmentation, A/B evaluation and more.

We're also currently testing supports for custom scripts – currently available for testing in the `"html"` interface. This lets you define your own actions and interfaces via custom recipes. See here for details and examples:

> [@Custom view templates with scripts](https://support.prodi.gy/t/custom-view-templates-with-scripts/302/6):
>
> Update: All works pretty smoothly already! :tada [custom\_js\_poc] The above example only needed the following html\_template and javascript config: \<button class="custom-button" onClick="updateText()"\> point_down Text to uppercase \</button\> \<br /\> \<strong\>{{text}}\</strong\> let upper = false; function updateText() { const text = window.prodigy.content.text; const newText = !upper ? text.toUpperCase() : text.toLowerCase(); window.prodigy.update({ text: newText }); upper = !upper; …
