TensorFlow dataset pipeline not reading some downloaded PPT/PDF files correctly

Hi everyone,

I am working on a small document dataset for a TensorFlow project. I have some presentation files in PPT/PDF format and I am trying to extract the text from them before putting the data into my TensorFlow pipeline.

Some of the presentations I downloaded from SlideShare using a SlideShare Downloader are working fine, but some files are giving me empty text or broken text after extraction.

I checked the files manually and they open normally in PowerPoint/PDF viewer. So I am not sure if this is a TensorFlow dataset issue or the way I am extracting the files before loading them.

My workflow is basically:

PPT/PDF → extract text → clean text → TensorFlow dataset → training

Has anyone worked with presentation documents in a TensorFlow dataset before? Is there any recommended way to preprocess these files so the extracted text is consistent?

I am mainly trying to understand where the problem is coming from, before changing my whole pipeline.

Any advice would be appreciated.