Sırdaş

PDF to Markdown for RAG

Prepare your PDF for RAG: clean Markdown, redacted personal data and a token count. All on your device.

Sent to servers: 0 BNothing left your device

About this tool

RAG answers better with well-structured text: headings, lists and paragraphs, without repeated headers or stray page numbers. We turn your PDF into Markdown with that structure and estimate the tokens so you know whether it fits in one conversation.

Before you paste anything into RAG, personal data is replaced with consistent placeholders ([NOMBRE_1], [DOCUMENTO_1]…): the same value always gets the same placeholder, so the model understands the document without knowing the people.

For RAG and embeddings, Markdown with clean headings and paragraphs yields coherent, retrievable chunks.

How it works

  1. 1

    Drop the PDF.

  2. 2

    Review the Markdown and the highlighted data.

  3. 3

    Copy (or download the .md) and paste it into RAG.

Questions

Can't RAG read the PDF directly?
Many assistants accept PDFs, but uploading sends the file to their servers. With Sırdaş the PDF never leaves your machine: you only share with RAG the anonymized text you choose to paste.
Which data is hidden before pasting into RAG?
ID and passport numbers, tax IDs, emails, phones, dates, addresses, cards, medical record numbers and labelled names (patient, landlord, signed by…). Detection is automatic and can miss items: review the result.
How do I know the token count?
We show an estimate (about 4 characters per token). It is indicative; each model counts slightly differently.

More tools