writing-style-extractionlisted
Install: claude install-skill popjam-io/skills
# Writing Style Extraction
Reverse-engineer a brand's verbal DNA from the copy it already publishes, and codify it so a copywriter or a generation pipeline reproduces *that* voice instead of a generic one.
The core insight mirrors visual guideline extraction: a voice extracted from N texts is only as good as (a) how representative the corpus is across channels, and (b) how rigorously you separate *rules* (what the brand always does) from *variations* (what it does per channel, campaign type or language) and *exceptions* (one-offs). A brand that writes 40-character emoji-led captions and 1,500-word how-to articles has one voice and two systems — averaging them produces a voice nobody wrote. The whole workflow is built around that separation.
Work through five phases in order. Phase 3's extraction fan-out is the expensive step; everything else is cheap.
## Phase 1 — Corpus inventory
Locate the texts. Usually the user provides a folder or an export (CSV/JSON of posts, a crawl, an ad-library dump); if they name sources (Instagram, blog, Meta Ad Library, website) gather what's accessible first, but never block on missing sources — work with what exists and record the gaps in the coverage section.
Build `work/corpus.jsonl`, one line per text unit, tagged by channel:
```json
{"id": "ig-2026-04-12", "channel": "social | long_form | ads | web", "text": "verbatim", "language": "tr",
"date": "2026-04-12", "engagement": {"likes": 412, "comments": 9}, "source_hint": "instagram | b