mdflux logo

mdflux

Turn any document into clean, AI-ready Markdown. Local-first desktop app: reads scanned PDFs, batches folders, runs offline, and uses far fewer tokens than vision models.

download v0.1.0platform Windowsvs vision up to 6× fewer tokensworks offline
Website GitHub

What is it?

What it is

A local, desktop application that turns PDFs, Office files, EPUBs, HTML, CSV, JSON, XML, images, and audio into clean, AI‑ready Markdown, adding OCR for scanned PDFs and optional cleanup passes.

Why it exists

To provide a privacy‑by‑default, offline solution that reduces token costs for LLM ingestion by delivering structured Markdown instead of sending image pages to vision models, while simplifying batch conversion of diverse document formats.

Who should use it

Teams building with Svelte who want an open-source, self-hosted option.

Who should avoid it

Teams that need a fully managed SaaS with enterprise SLAs out of the box.

How it works

A quick walkthrough in plain English

How mdflux works

Step 1 of 3

You interact with it

Open mdflux, send a request, or connect it to your stack.

Features

Fewer tokens, lower cost – Clean Markdown costs about 2 to 6 times fewer tokens than sending pages to a vision model
Local and private – Documents never leave your machine; no cloud, no API key, no account
Reads scanned PDFs – Built-in OCR recovers text that plain extractors return as zero characters
Real structure – Proper Markdown with headings, tables, and lists intact; readable, greppable, diff-able
No terminal needed – Portable app; unzip, run, click through a one-time setup
Many formats – Supports PDF, DOCX, PPTX, XLSX, EPUB, HTML, CSV, JSON, XML, images, and audio
Batch a whole folder – Convert everything at once with progress, cancellation, and per-file diagnostics
Optional cleanup – Off, rule-based, or AI pass (local or API) to tidy up messy extractions

Advantages

  • Reduces token usage by 2–6× compared to vision models, lowering LLM costs
  • Fully local and private; no data leaves your machine
  • Handles scanned PDFs via OCR, extracting text other tools miss
  • Produces clean, structured Markdown preserving headings, tables, lists
  • Portable desktop app with no terminal or command line required
  • Supports a wide range of file formats including Office, PDF, EPUB, HTML, CSV, JSON, XML, images, audio
  • Batch processing of entire folders with progress tracking and cancellation
  • Optional cleanup modes (off, rule-based, AI) to improve extraction quality
  • Works offline after initial setup
  • No account, API key, or cloud service needed
  • Dependency health and diagnostics panel for troubleshooting
  • Built on Microsoft's MarkItDown engine ensuring high-quality core conversion

Disadvantages

  • Currently available only for Windows 10/11 (x64)
  • First launch requires internet to set up the local Python environment
  • Unsigned build triggers Windows SmartScreen warning (requires 'Run anyway')
  • macOS support is planned but not yet available
  • No command-line interface (CLI) yet; only GUI app
  • Lacks code signing, which may cause security warnings
  • Requires WebView2 runtime (included with modern Windows but may need update)

Installation

native

1. Drop a file or folder   →   PDF, Office, EPUB, scans, audio, and more
2. Pick a cleanup mode      →   Off, rule-based, or AI (local or API)
3. Get clean Markdown       →   Preview, copy, or save as .md. 100% offline.

FAQ

What file formats does MDFlux support?

MDFlux supports PDF (including scanned PDFs via OCR), DOCX, PPTX, XLSX, EPUB, HTML, CSV, JSON, XML, images (PNG, JPG, GIF, WEBP, TIFF, BMP), and audio files (MP3, WAV, M4A, OGG, FLAC, AAC) for transcription.

How do I install and set up MDFlux?

Download the portable zip from the Releases page, extract it anywhere, and double-click MDFlux.exe. The first launch requires internet to set up a local Python environment (one-time setup). All subsequent conversions run fully offline.

Can MDFlux handle scanned PDFs that other tools can't read?

Yes, MDFlux includes built-in OCR that recovers text from scanned, image-only PDFs. While plain text extractors return zero characters, MDFlux's OCR recovers the full text and converts it to clean Markdown.

How much does MDFlux save on AI token costs?

MDFlux produces clean Markdown that costs 2 to 6 times fewer tokens than sending documents to vision models. For ordinary documents, it's about 4 times lighter, and up to 5.7 times lighter on scanned pages. This saving compounds across every LLM call that reads the document.

Does MDFlux work offline after initial setup?

Yes, after the one-time first launch setup (which requires internet), MDFlux runs entirely offline. Your documents never leave your machine, making it privacy-focused and suitable for sensitive materials.

What should I do if Windows shows a SmartScreen warning?

The portable build is unsigned but open source. Click 'More info' then 'Run anyway' to proceed. The WebView2 runtime is already included on current Windows 10/11 systems.

Loading documentation…
View on GitHub

Featured in Videos

YouTube tutorials and walkthroughs for mdflux

sd

80,246 views

Starts at 00:00

Share

Alternatives

Similar projects ranked by category, topics, and text overlap.

Compare
mdflux | MushyBook