Skip to content

MarkItDown (Microsoft doc→markdown converter) installed and wired into AI pipeline standards.

  • Installed WSL-only, not Windows: uv tool install 'markitdown[all]' → ~/.local/bin/markitdown. AI pipelines (Accounting, etc.) run WSL-side with uv; a Windows copy would be a divergent second install (SSOT). Talbot confirmed.
  • markitdown[all] (full extras: OCR/audio/onnx) chosen over a minimal install — covers PDF/DOCX/PPTX/XLSX/images/audio/HTML without a second install later.
  • Verified live against a real PDF — clean markdown output.
  • Core/IT/Utils/External/Utilities.md — new ## WSL section, MarkItDown reference + usage.
  • ai-config/AGENTS.md — new Dev Standards entry: agents should reach for markitdown instead of hand-rolling PDF/DOCX extraction. Deployed live (commit 7893432), verified in deployed ~/.claude/CLAUDE.md.
  • ai-config/claude/commands/process-expenses.md — Step 1 now points at markitdown for image-heavy bills, replacing a stale ~/utils/ai/pdf2md reference.
  • Ideas-Workflow/UPGRADES.md — deferred note: re-evaluate MarkItDown vs. the existing PyMuPDF-based pdf-to-clippings.py (custom Globe & Mail glyph-spacing fix) before extending that pipeline further.

~/utils/ai/pdf2md is a different, still-valid tool — AI-assisted (Gemini/Claude) extraction purpose-built for Globe & Mail newspaper articles with tag/frontmatter logic. Not a duplicate; just no longer the default suggestion for generic bill/document extraction.