The short version
I ran Microsoft’s converter against real files on one machine. Text formats and Office documents convert cleanly. Image-only PDFs do not convert at all — and the tool does not warn you.
What was measured
| Input | Result |
|---|---|
| HTML | Text and tables preserved |
| CSV | Converted to a Markdown table |
PowerPoint .pptx |
Slide by slide, titles kept |
Excel .xlsx |
Sheets and tables kept |
| PDF with a text layer | Text extracted correctly |
| PDF with no text layer | Exit code 0, 2 bytes of output |
The last row is the one that matters.
Why the PDF case is the dangerous one
A scanned PDF contains no font objects at all — only images. There is nothing for a text extractor to read. The tool still returns success, and you get an empty file that looks like a successful conversion.
You can verify this yourself. A PDF that has text carries /Font objects in its byte stream.
An image-only PDF carries /Image objects and zero fonts. That is the exact test this site
runs before it shows you a result.
What to do if your PDF has no text layer
You do not need a converter — you need OCR. Running it through the tool again will produce the same empty file. Check first, then choose the right tool.
Related
Convert a PDF, Word or Excel file to Markdown in the browser