Feature Request: Option to Disable Post-Processing and Preserve Detection Tags

#18
by UsmanTahirKiani - opened

Hi,

I'm using Unlimited-OCR for document parsing and have noticed that the generated Markdown output is already post-processed. It appears that functions such as remove_det() (or similar post-processing) strip the original detection tags before the results are written to the output file.

Currently, the output is plain Markdown/text, for example:

Introduction

This is the first paragraph...

However, I need access to the original structural predictions from the model, such as:

<|det|>title<|/det|>Introduction
<|det|>text<|/det|>This is the first paragraph...
<|det|>table<|/det|>| Col1 | Col2 |

Having these tags would allow me to distinguish titles, headings, body text, tables, captions, formulas, and other document elements so I can perform my own custom post-processing and generate Markdown according to my application's requirements.

Would it be possible to add an option such as:

model.infer_multi(
    ...,
    save_raw_output=True,      # Save output before remove_det()/post-processing
)

or

model.infer_multi(
    ...,
    post_process=False
)

This could save the raw decoder output (including the <|det|> tags) alongside the existing processed Markdown output.

This feature would be very useful for downstream applications that require document structure information instead of only flattened text.

Thank you!

This will be available in the native transformers implementation: https://github.com/huggingface/transformers/pull/46836

Sign up or log in to comment