# simulating prodigy user

**URL:** <https://support.prodi.gy/t/simulating-prodigy-user/5893>\
**Category:** Uncategorized\
**Tags:** usage\
**Created:** [August 29, 2022, 2:16pm UTC](https://support.prodi.gy/t/simulating-prodigy-user/5893 "2022-08-29T14:16:34Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [August 29, 2022, 9:21pm UTC](https://support.prodi.gy/t/simulating-prodigy-user/5893/2 "2022-08-29T21:21:46Z")

</div>

hi @FarisHijazi!

Thanks for your message and welcome to the Prodigy community 👋

Interesting project!

> [@FarisHijazi](#):
>
> I'm trying to run a grid search on an active learning pipeline, and I want to do this automatically without actually having a human in the loop, I need to somehow be able to simulate prodigy and connect to it's database without launching the server.

Curious - is your goal to run experiments to determine the best strategies on using active learning for improved accuracy? Interesting stuff! I haven't seen Robert's code but I would be interested to learn more. I can get back to you later.

One interesting thing for "simulating" active learning is to also add noise to the annotations (i.e., purposely make x% of annotations incorrect). If you do a simulation where you assume the annotator is correct each time, this isn't reflective on the reality that annotators make mistakes. I've devised AL experiments in the past and found an important "hyperparameter" is the assumed accuracy of the annotators. Just another factor to consider in your experiments.

> [@FarisHijazi](#):
>
> Here's the list of what I need to do:
> 
> - get access current annotations
> - get all previous annotations (I can do this by running the `db-out` and then parsing the output string in python), but I'm sure there's a better way
> - write current session annotations to prodigy db (only needed for the simulation to simulate a user labeling)

Have you [heard/used Prodigy's entry points](https://prodi.gy/docs/install#entry-points)?

> Entry points let you expose parts of a Python package you write to other Python packages. This lets one application easily customize the behavior of another, by exposing an entry point in its `setup.py` or `setup.cfg`. For a quick and fun intro to entry points in Python, check out [this excellent blog post](https://amir.rachum.com/blog/2017/07/28/python-entry-points/). Prodigy can load custom function from several different entry points, for example custom recipe functions. To see this in action, check out the [`sense2vec`](https://github.com/explosion/sense2vec) package, which provides several custom Prodigy recipes. The recipes are registered automatically if you install the package in the same environment as Prodigy. The following entry point groups are supported:

| | |
| --- | --- |
| `prodigy_recipes` | Entry points for recipe functions. |
| `prodigy_db` | Entry points for custom [`Database` classes](https://prodi.gy/docs/api-database). |
| `prodigy_loaders` | Entry points for custom [loader functions](https://prodi.gy/docs/api-loaders). |

I haven't used these yet but they may do the trick. Here's [where they were used in `sense2vec`](https://github.com/explosion/sense2vec/blob/master/setup.cfg#L37).

Let me think more in general -- I may have some suggestions.

In the meantime, I found a relevant post (which you may have read already):

> [@Is it possible for me to control the entire active learning loop?](https://support.prodi.gy/t/is-it-possible-for-me-to-control-the-entire-active-learning-loop/415/2):
>
> This is not necessarily true. Most use cases definitely involve loading data from a file or a single source, because that’s the most common way people go about annotating their data. But in the end, the stream is just a Python generator that Prodigy keeps requesting batches of tasks from. How that batch is composed is up to you – so you could easily implement your own logic that takes previous user decisions into account, randomly adds data from different sources or uses other factors to determ…

---

_[View the full topic](https://support.prodi.gy/t/simulating-prodigy-user/5893)._
