When you give a PDF to an AI chat or agent, many services use every page twice: they read the extracted text, and they also look at an image of the page so they can follow charts and layout. If the PDF is mostly words, you can skip the second part by converting it to Markdown first and pasting the text instead, whether you use ChatGPT, Claude, Gemini or another tool. Markdown is plain text, so every AI can read it. How many tokens that saves depends on how the service counts a PDF, and with some it saves nothing.

How AI services count a PDF

Three providers publish how their APIs handle PDFs. The chat apps built on them may work differently and mostly don’t say.

APIWhat each page becomesCost per page
OpenAI (GPT-4o and later)Extracted text and a page imageNot given as one number; the docs say page images “can increase token usage”
Anthropic ClaudeExtracted text and a page imageText about 1,500–3,000 tokens, plus the image (an A4 page at most about 1,551 tokens, or about 4,756 on Claude 4.7 and later, by the image formula)
Google GeminiPage read with vision; on Gemini 3 the PDF’s own text is not chargedGemini 3: about 560 tokens by default (280 low, 1,120 high resolution). Earlier models: 258 tokens

So converting helps most where text and page images are both counted. Anthropic gives one measured example, for Amazon Bedrock: a 3-page PDF used about 1,000 tokens with text extraction only and about 7,000 tokens with full visual understanding. On the Gemini API it can go the other way, because a dense page of pasted text can cost more than the 258 to 560 tokens a page costs there.

What we measured

We ran four public or self-made PDFs through PDF to Markdown with the default settings and counted what came out. We didn’t send them to any AI, so there are no measured token counts here. The last column uses Anthropic’s published image formula, the one service with a formula, to show the most the page images could have cost there.

PDFPages · sizeMarkdownClaude page images, at most
Research paper, single column (Attention Is All You Need)15 · 2.2 MB39,909 characters (40 KB), 6 tablesabout 22,000 tokens, or 71,000 on Claude 4.7+
Research paper, two columns (Deep Residual Learning, ResNet)12 · 819 KB62,294 characters (63 KB)about 18,000 tokens, or 57,000 on Claude 4.7+
Tax form (IRS Form W-4, 2026)5 · 209 KB27,926 characters (28 KB), 12 tablesabout 7,500 tokens, or 24,000 on Claude 4.7+
Korean guide saved as PDF3 · 123 KB3,179 characters (5.7 KB), 4 tablesabout 4,700 tokens, or 14,000 on Claude 4.7+

The paper’s Markdown is under 2% of the PDF’s file size, but no AI service charges by file size; the point is that all the text is there in a much smaller package. A page with little text saves the most in proportion where page images are charged, because its image costs as much as a full page.

When converting helps

  • Text-heavy documents: reports, contracts, terms, papers, manuals, meeting notes.
  • Long documents: the more pages, the more page images you avoid. Some apps already switch to text only for long files; claude.ai, for example, analyzes visuals only in PDFs of 100 pages or fewer.
  • Long chats and agents: the document is part of every turn, so a smaller version keeps each request lighter.
  • Tools that don’t take PDFs: plain text works anywhere, including local models and older APIs.

When to keep the PDF

  • Charts, graphs and diagrams carry information the text doesn’t. Converting drops them.
  • Scans and photos of pages have no text layer, so there is nothing to convert. The tool tells you which pages came out empty.
  • Layout-heavy pages such as forms with checkboxes, slides, or tables with merged cells can come out shifted. Check the preview before you rely on a table.
  • Services that bill a PDF page at a flat rate, such as the Gemini API, where the PDF can already be the cheaper option.

How to do it

  1. Open PDF to Markdown and drop in the PDF. It is read on your device; nothing is uploaded.
  2. Leave the page markers on if you want to ask about specific pages, and leave repeated headers and footers removed.
  3. Press Convert to Markdown, read the preview, then Copy text or download the .md file.
  4. Paste it into your AI chat (or attach the .md file) and ask your question. If one chart matters, attach just that page as an image alongside the text.

On the command line: MarkItDown

Microsoft’s free, open-source MarkItDown does the same job from a terminal and also handles Word, PowerPoint and Excel files. With Python installed, pip install 'markitdown[all]' installs it and markitdown report.pdf -o report.md converts a file. It is a good fit for converting many files at once; the page here needs no install and keeps headings, lists and tables from the page layout.

Sources