# Create baseline metrics based on manual NER annotations

**URL:** <https://support.prodi.gy/t/create-baseline-metrics-based-on-manual-ner-annotations/2998>\
**Category:** Uncategorized\
**Tags:** usage, ner, solved\
**Created:** [June 8, 2020, 9:44am UTC](https://support.prodi.gy/t/create-baseline-metrics-based-on-manual-ner-annotations/2998 "2020-06-08T09:44:59Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![svenski](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/svenski/32/1459_2.png) [@svenski](https://support.prodi.gy/u/svenski)\
**Post date:** [June 8, 2020, 9:44am UTC](https://support.prodi.gy/t/create-baseline-metrics-based-on-manual-ner-annotations/2998/1 "2020-06-08T09:44:59Z")

</div>

I'm taking my first steps with Prodigy and have annotated a test set with ORG labels (only).

My intention is to see how well the existing Spacy models, and eventually other NER models, perform on this labelled data set out of the box, before I start any training. I haven't been able to find a simple way to do this. I essence I'm looking for a _evaluate.model_ recipe where I can pass in a test/validation set and a model which outputs evaluation metrics.

I tried passing in no training data to the _train_ recipe but it didn't want to play ball.

Currently I'm trying to transform the output from the _data-to-spacy_ recipe to fit with the _GoldParse_ input, so I can get some . It seems to me like it could be standard use case, so wanted to check if there is a simpler way to achieve what I want?

---

<div class="post-metadata">

**Author:** ![svenski](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/svenski/32/1459_2.png) [@svenski](https://support.prodi.gy/u/svenski)\
**Post date:** [June 8, 2020, 12:19pm UTC](https://support.prodi.gy/t/create-baseline-metrics-based-on-manual-ner-annotations/2998/2 "2020-06-08T12:19:11Z")

</div>

It seems like the following post is dealing with the same question:

> [@Evaluating Precision and Recall of NER](https://support.prodi.gy/t/evaluating-precision-and-recall-of-ner/193/4):
>
> In case others like me come looking for a basic scoring recipe, here is what I cooked up. It doesn’t consider threshold, but it evaluates model accuracy without re-training and can output either PRF or the standard prodigy score scheme. import spacy import spacy.scorer from prodigy.components.db import connect from prodigy.core import recipe, recipe\_args from prodigy.models.ner import EntityRecognizer, merge\_spans from prodigy.util import log from prodigy.components.preprocess import split\_sen…

I'm trying this now.

---

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [June 8, 2020, 12:21pm UTC](https://support.prodi.gy/t/create-baseline-metrics-based-on-manual-ner-annotations/2998/3 "2020-06-08T12:21:07Z")

</div>

Hi! I think what you're looking for is `spacy evaluate`?

> **[Command Line Interface · spaCy API Documentation](https://spacy.io/api/cli/)**
>
> Download, train and package models, and debug spaCy

This takes data in spaCy's format and will perform an evaluation. Prodigy's training experiments are really designed for quick training experiments with Prodigy dataset and not necessarily to replace the training or evaluation process of whichever library you're using (e.g. spaCy).

The above post is a bit more abstract and about a custom evaluatio, including evaluation of binary data (I think) so I'm not sure that's the right approach if you just want to output scores.

---

<div class="post-metadata">

**Author:** ![svenski](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/svenski/32/1459_2.png) [@svenski](https://support.prodi.gy/u/svenski)\
**Post date:** [June 8, 2020, 12:41pm UTC](https://support.prodi.gy/t/create-baseline-metrics-based-on-manual-ner-annotations/2998/4 "2020-06-08T12:41:53Z")

</div>

Thank you for the prompt reply I didn't know about `spacy evaluate`! However, I will log the results to a database so using the code from the example above fits the bill quite well.
