defuddlelisted
Install: claude install-skill tboome33/obsidian-mcp-router
# defuddle
Cheap content cleaning before ingestion. Worth the extra step on most webpages — almost free in tokens after, much cheaper to ingest, and the wiki page that comes out is more readable.
## When to use
- Before `wiki-ingest` on a URL that's a typical webpage (blog post, news article, documentation site).
- Inside `autoresearch` for any HTML fetch.
- The user pastes a URL and says "what does this say" — defuddle it then summarize.
## When NOT to use
- The URL is already a clean source (raw markdown, GitHub raw, RSS feed, JSON API).
- The URL is a PDF or video — defuddle is for HTML.
- The user wants the raw HTML for some specific reason.
## Steps
### 1. Fetch the page
Use `WebFetch` with a prompt like:
> Return the main article content as clean markdown. Drop navigation, ads, cookie banners, related posts, comment sections, social media widgets, footer boilerplate, and "subscribe to our newsletter" callouts. Preserve: the article title, author, publication date if visible, headings, body paragraphs, code blocks, lists, blockquotes, inline links to relevant resources (drop tracking-only links), images that are part of the content (note their captions if any).
WebFetch's underlying small model is good at this kind of selective extraction. Trust its output.
### 2. Validate the output
Quick sanity checks:
- Length: if the cleaned output is < 200 chars or > 50K chars, something probably went sideways. Surface to the user with the URL and offer to retry or fall b