# Losing spancat labels when training after using prodigy db-merge

**URL:** <https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995>\
**Category:** Uncategorized\
**Tags:** spacy, spancat\
**Created:** [December 15, 2023, 2:17pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995 "2023-12-15T14:17:04Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 15, 2023, 2:17pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/1 "2023-12-15T14:17:04Z")

</div>

Hello there. These are great applications, thank you for all of the effort and support.

I'm creating spancats on multiple types of documents, but looking for the same labels across them. For both ease of annotation and to see if it would make much of a difference, I've trained each one separately. When I evaluate each one independently, I'm getting good R,P, and F scores, but when I merge them, I'm losing entire labels. What's really odd, is that both of the datasets have annotations for "PATIENT" and "DOB" - if I was losing labels that only appear in one (e.g. ENCDATE) I would, maybe, understand. I have 27 different "databases" -- but to simplify what I'm seeing, I've reproduced it with 2.

My evaluation of the first "allergy" documents:

 ![Screenshot 2023-12-15 at 9.04.23 AM](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/b/bc4c20f3e9da347e350c3d5f605d31e47c822a48.png)

My evaluation of the second "audiology" documents:

 ![Screenshot 2023-12-15 at 9.04.58 AM](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/7/7d6b42bf90eb4be057a3e6a3a7537295f34148a1.png)

The command I'm using to merge the data:  
prodigy db-merge FAX\_ALLERGY,FAX\_AUDIOLOGY FAX\_MERGED\_AUDIOLOGY\_ALLERGY

Q: Does the order matter?

I'm then converting to spacy:  
python -m prodigy data-to-spacy spacy --spancat FAX\_MERGED\_AUDIOLOGY\_ALLERGY --eval-split 0.2 --verbose

Then training:  
python -m spacy train spacy/config.cfg --output ./model --paths.train spacy/train.spacy --paths.train spacy/train.spacy --paths.dev spacy/dev.spacy --gpu-id 0 --verbose

Then evaluating:  
python -m spacy evaluate ./model/model-best ./spacy/dev.spacy --spans-key PATIENT,DOB,MRN,SSN,INSPOLN,INSGRPN,EMAIL,PHONE,ENCDATE,CLAIMN -o ./evaluation.json

The output of evaluation.json on the combined model:

 ![Screenshot 2023-12-15 at 9.05.33 AM](https://us1.discourse-cdn.com/flex020/uploads/prodigy/original/2X/d/d67cfe2dba484ca0b6e89e47dd6add6a1dba4b01.png)

I'm losing labels -- what am I doing incorrectly (or simply don't understand)?

Note: When I combined all 27 databases, I lose fewer labels, but I'm still losing some, which is less than ideal.

Thanks in advance! I'm sure this is PEBKAC, but I've been playing around for a few days and I'm stuck.

---

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 15, 2023, 2:32pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/2 "2023-12-15T14:32:11Z")

</div>

HOAS. I might just be an idiot... Of course, I realized this AFTER I posted.. I checked my config.cfg  
and noticed the label reader was pointing to a non-existent file. Very odd that it would find any labels at all. Let's see if the corrected labels fix things. Standby.

[initialize.components.textcat.labels]  
@readers = "spacy.read\_labels.v1"  
path = "spacy/labels/spancat.json"  
require = true

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [December 15, 2023, 2:37pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/3 "2023-12-15T14:37:46Z")

</div>

Thanks for the update.

Just to make sure, have you seen these recent posts?

> [@Disappearing spans when using data-to-spacy](https://support.prodi.gy/t/disappearing-spans-when-using-data-to-spacy/6992/4):
>
> As an aside for extra context all base English models inside of spaCy use the same tokeniser under the hood. So the tokens from nlp = spacy.blank("en") should be the same as those from spacy.load("en\_core\_web\_sm"). These tokens are all determined by the same rule based system. But yeah, it does sound like there's a mismatch. One avenue to explore is to retokenize everything by using [this spaCy method](https://spacy.io/api/span#char_span) which comes with a alignment\_mode parameter that should allow you to wiggle around minor charac…

> [@data-to-spacy losing annotations](https://support.prodi.gy/t/data-to-spacy-losing-annotations/6864/7):
>
> Thank you! I had to modify merge function to deal with the fact that not every example has a span for a label: def merge(examples): key\_values = {} for ex in examples: key = ex['\_input\_hash'] if key in key\_values: if 'spans' in ex: if 'spans' in key\_values[key]: key\_values[key]['spans'].extend(ex['spans']) else: key\_values[key]['spans'] = ex['spans'] else: key\_values[k…

One other debugging tip - did you know you can view `db-merge` or `data-to-spacy` recipes (or any other built-in recipes)? This way you can see exactly what they're doing and debug/modify them.

Just run `prodigy stats`, then look in the folder with your `Location:`, and find the `recipes` folder. For `db-merge`, it'll be in the `commands.py` script and `data-to-spacy` is in the `train.py`.

Hope this helps!

---

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 15, 2023, 5:54pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/4 "2023-12-15T17:54:12Z")

</div>

> [@dcane](#):
>
> [initialize.components.textcat.labels]

Sadly textcat labels and spancat aren't the same thing. I think it's far more likely that i need to edit the merge function of either or both db-merge and/or data-to-spacy.

---

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 15, 2023, 8:17pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/5 "2023-12-15T20:17:26Z")

</div>

Ok, after a day of attempts, I can confirm that the "patch" to the db-merge and prodigy-to-spacy creates train.spacy and dev.spacy docbins that have spans which include all of my labels.

```python
allspans_train = {}
for doc in docs_train:
    for span in doc.spans["sc"]:                
        if span.label_ not in allspans_train:
            allspans_train[span.label_]=1        
        else:
            allspans_train[span.label_]= int(allspans_train[span.label_]) + 1        
    
allspans_train 

{'ENCDATE': 695,
 'PATIENT': 3550,
 'DOB': 1791,
 'INSPOLN': 682,
 'PHONE': 251,
 'INSGRPN': 358,
 'MRN': 501,
 'SSN': 44,
 'EMAIL': 17,
 'CLAIMN': 78}

```

And the testing data:

```python
allspans_dev = {}
for doc in docs_dev:
    for span in doc.spans["sc"]:                
        if span.label_ not in allspans_dev:
            allspans_dev[span.label_]=1        
        else:
            allspans_dev[span.label_]= int(allspans_dev[span.label_]) + 1        
    
allspans_dev 

{'PATIENT': 809,
 'MRN': 93,
 'DOB': 421,
 'ENCDATE': 166,
 'PHONE': 50,
 'SSN': 5,
 'INSPOLN': 148,
 'INSGRPN': 72,
 'EMAIL': 4,
 'CLAIMN': 18}

```

So - we are now getting all of the labels in the dataset merged.  
BUT -- still no dice on training. (Tried using prodigy and on the exported spacy files that I verified contains the spans)

python -m spacy evaluate ./spacy\_model/model-best ./spacy/dev.spacy --gpu-id 0 --spans-key PATIENT,DOB,MRN,SSN,INSPOLN,INSGRPN,EMAIL,PHONE,ENCDATE,CLAIMN -o ./evaluation.json

```python
{
  "token_acc":1.0,
  "token_p":1.0,
  "token_r":1.0,
  "token_f":1.0,
  "spans_sc_p":0.7686628384,
  "spans_sc_r":0.5246360582,
  "spans_sc_f":0.6236272879,
  "spans_sc_per_type":{
    "DOB":{
      "p":0.8625954198,
      "r":0.8052256532,
      "f":0.8329238329
    },
    "MRN":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    },
    "PHONE":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    },
    "PATIENT":{
      "p":0.7239709443,
      "r":0.739184178,
      "f":0.7314984709
    },
    "ENCDATE":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    },
    "SSN":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    },
    "INSPOLN":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    },
    "INSGRPN":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    },
    "EMAIL":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    },
    "CLAIMN":{
      "p":0.0,
      "r":0.0,
      "f":0.0
    }
  },
  "speed":37372.5665677684
}

```

Help 🙂

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [December 15, 2023, 8:25pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/6 "2023-12-15T20:25:18Z")

</div>

> [@dcane](#):
>
> BUT -- still no dice on training. (Tried using prodigy and on the exported spacy files that I verified contains the spans)

Could you run `data debug`?

This post seems relevant:

> **[Unable to train Spancat model, accuracy stuck at 0.0 · explosion/spaCy ·...](https://github.com/explosion/spaCy/discussions/11423#discussioncomment-3523169)**
>
> Losses decrease to 0 within 2000-3000 epochs, f, p, r scores do not go higher than 0 (image 1). Then I tried a different datasets and got a result like this in image2. But when I tried the old data...

If you don't see anything in `data debug`, any chance you have?

```python
[training.score_weights]
spans_sc_f = 1.0
spans_sc_p = 0.0
spans_sc_r = 0.0

```

There are a few more related posts on the spaCy GitHub discussions. Since you're problem is more spaCy than Prodigy, you might want to check that repo for help.

Hope this helps!

---

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 17, 2023, 2:08am UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/8 "2023-12-17T02:08:48Z")

</div>

24 hours of trying all sorts of things later, I'm still stuck. I've tried a few dozen ways of training (CPU, GPU, efficiency, accuracy, etc) and I'm still losing entire labels. I'm parsing through the spacy discussions, but so far have not come across anything similar. I've also reproduced the same results on 3 different machines with 3 different GPUs just to see if it was something environmental. Any ideas?

---

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [December 18, 2023, 1:12pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/9 "2023-12-18T13:12:33Z")

</div>

hi @dcane,

Sorry you're having issues, especially with so much work. As I mentioned in my last post, can you check your what your config file provides for your `training.core_weights` like:

```python
[training.score_weights]
spans_sc_f = 1.0
spans_sc_p = 0.0
spans_sc_r = 0.0

```

Like the thread mentioned:

> It's possible to have zero scores in training but actually train a model if your logging is configured incorrectly.

Since it seems like your issue is more spaCy, would it be possible to post on the spaCy GitHub discussion? Be sure to post the entire `config` file as that will help debug the problem. That forum includes spaCy core developers who can help you a lot faster since it's seemingly like your core issue is spaCy, not Prodigy.

---

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 18, 2023, 3:47pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/10 "2023-12-18T15:47:43Z")

</div>

> [@ryanwesslen](#):
>
> [Unable to train Spancat model, accuracy stuck at 0.0 · explosion/spaCy · Discussion #11423 · GitHub](https://github.com/explosion/spaCy/discussions/11423#discussioncomment-3523169)

Good idea. FWIW - I'm not upset at all. This is learning 🙂 I did confirm that BOTH the prodigy-auto generated config.cfg (which happens when you **python -m prodigy data-to-spacy spacy --spancat FAX\_MERGED\_AUDIOLOGY\_ALLERGY --eval-split 0.2 --verbose** ) looks correct. I've run a diff against that and a blank spancat config that I populated with ( **python -m spacy init fill-config base\_config.cfg config.cfg** ), and BOTH of them have:

```python
[training.score_weights]
spans_sc_f = 1.0
spans_sc_p = 0.0
spans_sc_r = 0.0

```

The only differences between them are that the spacy config generated from prodigy has this at the bottom:

```python

[initialize.components.spancat]

[initialize.components.spancat.labels]
@readers = "spacy.read_labels.v1"
path = "spacy/labels/spancat.json"

```

And fill-config from space defaults n-grames to:

```python
[components.spancat.suggester]
@misc = "spacy.ngram_suggester.v1"
sizes = [1,2,3]

```

Whereas Prodigy uses:

```python
[components.spancat.suggester]
@misc = "spacy.ngram_range_suggester.v1"
min_size = 1
max_size = 9

```

But that shouldn't make much of a difference. I'm running one complete end-to-end test to verify that all of the steps I've taken are correct, and then I'm going to post my findings on the spacy discussion forum.

---

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 18, 2023, 6:55pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/11 "2023-12-18T18:55:09Z")

</div>

Something strange is going on. When I train using the .cfg from prodigy vs. the spacy one I'm getting different results. The prodigy-based .cfg file is losing only 1 label, while the spacy filled-in init is losing 2 labels. Same exact spacy docbins. Both are losing the "MRN" (which is in both datasets), but the spacy filled in config is also losing ENCDATE (which is only in 1 dataset). I've triple checked the docbins, and the labels are 100% there:

train.spacy

```python
{'ENCDATE': 692,
 'PATIENT': 2454,
 'DOB': 1588,
 'EMAIL': 19,
 'MRN': 491,
 'PHONE': 220,
 'INSPOLN': 94,
 'INSGRPN': 68,
 'SSN': 37,
 'CLAIMN': 3}

```

dev.spacy

```python
{'PATIENT': 543,
 'ENCDATE': 162,
 'DOB': 360,
 'MRN': 95,
 'SSN': 7,
 'PHONE': 60,
 'INSPOLN': 28,
 'INSGRPN': 14,
 'EMAIL': 3}

```

---

<div class="post-metadata">

**Author:** ![adriane](https://avatars.discourse-cdn.com/v4/letter/a/46a35a/32.png) [@adriane](https://support.prodi.gy/u/adriane)\
**Post date:** [December 19, 2023, 7:40am UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/12 "2023-12-19T07:40:02Z")

</div>

The most common span tokens above in `debug data` look unexpected, especially having email addresses be frequent enough to show up in this list, so maybe there is still some issue with the underlying data/annotation/merging? (Or maybe it's just related to some anonymization?)

For email, you should consider using a pattern with `LIKE_EMAIL`, and for SSN or phone numbers patterns may also work better in addition to or instead of `spancat`.

The ngram lengths could make a difference if there are entity types like phone numbers that are typically always 4-grams or longer. You can try the exact same suggester settings from prodigy to see if that makes a difference? (Prodigy looks at your annotation in order to pick reasonable ngram lengths for the suggester, but spacy just uses 1-3 by default without analyzing your annotation.)

---

<div class="post-metadata">

**Author:** ![dcane](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@dcane](https://support.prodi.gy/u/dcane)\
**Post date:** [December 19, 2023, 9:33pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/13 "2023-12-19T21:33:06Z")

</div>

Lots and lots more testing...

I'm starting to wonder - is this a "feature" of NER - not a problem with prodigy/spacy? So, let's say I have 500 annotated faxes of "allergy" - out of 500, let's say most if not all have a PATIENT and DOB, some have an MRN, but very few have ENCDATE, SSN, PHONE, etc.

When I train allergy by itself, it only can recall the first 3 things it's seen the most: PATIENT, DOB, AND MRN

When I look at audiology, we almost always have an ENCDATE, PATIENT, DOB, and MRN, but little else and the training evaluation reflects that.

When I MERGE both sets together, and then train, my scores where the label counts roughly overlap in frequency are similar, but I'm losing labels that have low label counts across the newly combined set.

/me waits for someone who actually knows what he/she is doing for a slap across the head

So, even though in aggregate I have more examples of each of the lesser annotated fields, the overall % of total examples across the dataset is low, and those labels are dropping. Expected?

---

<div class="post-metadata">

**Author:** ![magdaaniol](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/magdaaniol/32/2787_2.png) [@magdaaniol](https://support.prodi.gy/u/magdaaniol)\
**Post date:** [January 3, 2024, 12:04pm UTC](https://support.prodi.gy/t/losing-spancat-labels-when-training-after-using-prodigy-db-merge/6995/14 "2024-01-03T12:04:01Z")

</div>

Hi @dcane,

It might be the case the the model didn't learn the dropped categories well enough to generalize on the dev set.  
Have you tried what happens of you test on the train dataset? If you can recall all labels while testing on the train set, at least you'd know there's nothing structurally wrong with the data. If you can confirm that, the next thing I'd do is to try to upsample the underrpresented categories other than emails, SSNs and phone numbers. For emails, SSNs and phone numbers I'd follow @adriane 's advice on using patterns rather than `spancat`. Let us kow how it goes!
