# Page Classification of PDF Documents

**URL:** <https://support.prodi.gy/t/page-classification-of-pdf-documents/1118>\
**Category:** Uncategorized\
**Tags:** usage, custom\
**Created:** [January 12, 2019, 8:44pm UTC](https://support.prodi.gy/t/page-classification-of-pdf-documents/1118 "2019-01-12T20:44:18Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![glander](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@glander](https://support.prodi.gy/u/glander)\
**Post date:** [January 12, 2019, 8:44pm UTC](https://support.prodi.gy/t/page-classification-of-pdf-documents/1118/1 "2019-01-12T20:44:19Z")

</div>

Hello Prodigy Community,

I’m considering purchasing Prodigy for a client that would like to classify pages of pdf documents into about 30 classes. Model would be operating on string content of pages w/ some other derived features to help it (ie, pixel saturation by page section). They have plenty of training data and a reliable number of labels for half of those classes, the remainder of which will have to be manually annotated by some interns. I’m not sure how good a fit prodigy is for this and am trying to evaluate.

I can use mupdf and pymupdf to turn pages into images and do the classification through prodigy’s image classification, but am not sure how much customizability that has from the _Multiple Choise (Image)_ live demo. But it seems like a lot of work to implement as I’ll have to not only create a pipeline for forming the images, but one for tracking which strings corresponded to those annotations when I’m feeding back to my tf model.

So I guess the question I’m asking is, is my project the right use case for prodigy, or will I be better off just making my own ad hoc annotation client in something like pysimplegui?

Thank you for your input and I apologize if this question has been asked before, the closest I could find were the two threads linked below and neither was able to answer my question.

[Web page classification](https://support.prodi.gy/t/text-classification-content-of-a-web-page/785)  
[Prodigy with PDFs](https://support.prodi.gy/t/using-prodigy-with-pdf-documents/322)

---

<div class="post-metadata">

**Author:** ![honnibal](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/honnibal/32/35_2.png) [@honnibal](https://support.prodi.gy/u/honnibal)\
**Post date:** [January 14, 2019, 4:39pm UTC](https://support.prodi.gy/t/page-classification-of-pdf-documents/1118/2 "2019-01-14T16:39:46Z")

</div>

Hi @glander,

In general I often find myself encouraging people to consider self-implemented alternatives. We’re definitely of the belief that one software package can’t do everything — if it could implement any solution, Prodigy would be a language, not a library!

I think what you’re trying to do is a bit outside the core use-case, so it’s possible it’ll be better to start fresh. That said, Prodigy’s pretty easy to customise, so if you want a web-based UI, I think just the management of the database/REST/CLI interaction that Prodigy provides might be useful enough to make it worth working with.

If you want to try out the back-end, I could set up a demo VM for you — send us an email at contact@explosion.ai

Best,  
Matt
