Hi everyone,
I am working on a small document dataset for a TensorFlow project. I have some presentation files in PPT/PDF format and I am trying to extract the text from them before putting the data into my TensorFlow pipeline.
Some of the presentations I downloaded from SlideShare using a SlideShare Downloader are working fine, but some files are giving me empty text or broken text after extraction.
I checked the files manually and they open normally in PowerPoint/PDF viewer. So I am not sure if this is a TensorFlow dataset issue or the way I am extracting the files before loading them.
My workflow is basically:
PPT/PDF → extract text → clean text → TensorFlow dataset → training
Has anyone worked with presentation documents in a TensorFlow dataset before? Is there any recommended way to preprocess these files so the extracted text is consistent?
I am mainly trying to understand where the problem is coming from, before changing my whole pipeline.
Any advice would be appreciated.