Technology Aug 26, 2026 · 3 min read

Building Local-First Web Apps: Parsing HTML and PDFs to Markdown in the Browser

Local-first and privacy-focused web utilities are having a massive comeback. With browser engines becoming faster and WebAssembly/Web Workers maturing, there is rarely a reason to push sensitive user documents to an external backend for simple conversions. While building MD-Convert (a zero-upload d...

DE
DEV Community
by moamen abou elsaoud
Building Local-First Web Apps: Parsing HTML and PDFs to Markdown in the Browser

Local-first and privacy-focused web utilities are having a massive comeback. With browser engines becoming faster and WebAssembly/Web Workers maturing, there is rarely a reason to push sensitive user documents to an external backend for simple conversions.

While building MD-Convert (a zero-upload document to Markdown converter), I explored how to parse real-world documents into clean Markdown entirely on the client side.

Here is a breakdown of the core architecture and libraries that make purely in-browser document processing possible.

1. Converting Web Articles with Readability + Turndown

Converting messy web markup into clean Markdown involves two distinct steps:

  1. Content Extraction: Stripping ads, navbars, sidebars, and trackers.
  2. HTML-to-Markdown Transformation: Translating semantic DOM nodes into markdown tokens.

Mozilla’s @mozilla/readability paired with turndown is an incredible combination for this:


javascript
import { Readability } from '@mozilla/readability';
import TurndownService from 'turndown';

function htmlToCleanMarkdown(rawHtmlDocument, sourceUrl) {
  // 1. Extract pure article content
  const reader = new Readability(rawHtmlDocument);
  const article = reader.parse();

  if (!article || !article.content) {
    throw new Error('Unable to extract main content');
  }

  // 2. Initialize Turndown
  const turndownService = new TurndownService({
    headingStyle: 'atx',
    codeBlockStyle: 'fenced'
  });

  // Ensure image URLs remain absolute
  turndownService.addRule('absoluteImages', {
    filter: 'img',
    replacement: (content, node) => {
      const src = node.getAttribute('src');
      const alt = node.getAttribute('alt') || '';
      if (!src) return '';
      try {
        const absoluteUrl = new URL(src, sourceUrl).href;
        return `![${alt}](${absoluteUrl})\n\n`;
      } catch {
        return `![${alt}](${src})\n\n`;
      }
    }
  });

  return turndownService.turndown(article.content);
}

2. Offloading Heavy PDF Parsing to Web Workers
Parsing large PDFs using pdf.js on the main thread is a recipe for UI freezes and dropped frames. The solution is running the extraction pipeline inside a dedicated Web Worker.
Here is a simplified pattern for extracting selectable text layers asynchronously:
// worker.js
import * as pdfjsLib from 'pdfjs-dist/legacy/build/pdf';

self.onmessage = async (e) => {
  const { arrayBuffer } = e.data;

  try {
    const loadingTask = pdfjsLib.getDocument({ data: arrayBuffer });
    const pdf = await loadingTask.promise;
    let fullText = '';

    for (let pageNum = 1; pageNum <= pdf.numPages; pageNum++) {
      const page = await pdf.getPage(pageNum);
      const textContent = await page.getTextContent();

      const pageText = textContent.items
        .map((item) => item.str)
        .join(' ');

      fullText += `## Page ${pageNum}\n\n${pageText}\n\n`;
    }

    self.postMessage({ status: 'success', markdown: fullText });
  } catch (error) {
    self.postMessage({ status: 'error', message: error.message });
  }
};

The Big Advantages of Client-Side Processing
Zero Server Costs: The application can run entirely on static hosting (like Cloudflare Pages or GitHub Pages).
Absolute Privacy: User data, API keys, or confidential spreadsheets never cross the network.
Instant Latency: Conversions happen in-memory without queue waiting times.
Thoughts & Edge Cases?
The main challenge with 100% in-browser parsing remains edge cases—like complex multi-column PDF layouts or non-OCR scanned documents.
Have you built any client-side/local-first tools recently? What libraries do you prefer for client-side parsing? Let's discuss in the comments! 👇
DE
Source

This article was originally published by DEV Community and written by moamen abou elsaoud.

Read original article on DEV Community
Back to Discover

Reading List