# Can we train abstractive summarization in prodigy?

**URL:** <https://support.prodi.gy/t/can-we-train-abstractive-summarization-in-prodigy/6475>\
**Category:** Uncategorized\
**Tags:** usage, textcat\
**Created:** [April 3, 2023, 1:08pm UTC](https://support.prodi.gy/t/can-we-train-abstractive-summarization-in-prodigy/6475 "2023-04-03T13:08:16Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ryanwesslen](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ryanwesslen/32/2969_2.png) [@ryanwesslen](https://support.prodi.gy/u/ryanwesslen)\
**Post date:** [April 3, 2023, 1:43pm UTC](https://support.prodi.gy/t/can-we-train-abstractive-summarization-in-prodigy/6475/2 "2023-04-03T13:43:00Z")

</div>

hi @mirfan923!

Thanks for your question. This looks like an interesting repo.

> [@mirfan923](#):
>
> Can we train the abstractive summarization on these datasets using prodigy?

It's important to remember that Prodigy is an annotation tool to acquire more annotated data, not necessarily a tool for model training.

Prodigy does have the[`train` recipe](https://prodi.gy/docs/recipes/#training), but it's just a wrapper for `spacy train`. Since spaCy doesn't haven't a built-in summarization component, it's not possible to train abstractive summarization out-of-the-box with `prodigy train`.

Since you mentioned the "datasets" in the repo - are you only interested in training or using the datasets and model in the repo, and creating a "model-in-the-loop" workflow to acquire more annotated data?

If you're only interested in training with those data and not getting any additional annotated data, then I'm not sure Prodigy would help.

However, if you wanted a model-in-the-loop workflow, then yes, Prodigy could help if you wrote a [custom recipe](https://prodi.gy/docs/custom-recipes). Custom recipes are essentially Python functions (written as a Python script) that can be run through the command line. So in this way, it may be possible you could write a custom recipe to do abstractive summarization with another model framework (e.g., the [`seq2seq` training module](https://github.com/csebuetnlp/xl-sum/tree/master/seq2seq) used in the repo you posted).

Alternatively, if you only wanted additional annotated data for summarization (no model in the loop), you could create a custom recipe like this:

> [@Labelling dataset for extractive text summarization](https://support.prodi.gy/t/labelling-dataset-for-extractive-text-summarization/3500/2):
>
> I was actually able to figure out Q1, but still unsure about Q2 and Q3. Here is the recipe.py file: import prodigy from prodigy.components.loaders import JSONL @prodigy.recipe( "extsumm", dataset=("The dataset to save to", "positional", None, str), file\_path=("Path to texts", "positional", None, str), ) def extsumm(dataset, file\_path): """Annotate sentences of a document to be included in extractive summary or not.""" stream = JSONL(file\_path) # load in the JSONL file …

You could also create a [custom interface](https://prodi.gy/docs/custom-interfaces) based on what annotation task you were looking for:

> [@Extractive summarization with labels](https://support.prodi.gy/t/extractive-summarization-with-labels/5702/2):
>
> I'm unaware of a Prodigy interface that offers this functionality out of the box. So it sounds like you might be interested in designing [a custom interface](https://prodi.gy/docs/custom-interfaces) for your specific task. Given the high level of interaction you require, especially with the editable summary, this could be a lot of work. So while the custom interface could be a valid option, I wonder if it's possible to simplify your interface instead. It sounds like one part of the problem is selecting sentences, which is something that…

Hope this helps!

---

_[View the full topic](https://support.prodi.gy/t/can-we-train-abstractive-summarization-in-prodigy/6475)._
